EDBT 2026 Demo / reviewers in the wild / expert
Rui Yan 0010
dblp:19/2405-10
· DBLP profile ↗
45ranked-venue papers
10as first author
41since 2021 · last 2026
0000-0002-0694-9458ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 6 first-author · 32 since 2021Artificial intelligence and machine learning · 17 · 8 first-author · 15 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spatiotemporal-Untrammelled Mixture of Experts for Multi-Person Motion PredictionabstractComprehensively and flexibly capturing the complex spatio-temporal dependencies of human motion is critical for multi-person motion prediction. Existing methods grapple with two primary limitations: i) Inflexible spatiotemporal representation due to reliance on positional encodings for capturing spatiotemporal information. ii) High computational costs stemming from the quadratic time complexity of conventional attention mechanisms. To overcome these limitations, we propose the Spatiotemporal-Untrammelled Mixture of Experts (ST-MoE), which flexibly explores complex spatio-temporal dependencies in human motion and significantly reduces computational cost. To adaptively mine complex spatio-temporal patterns from human motion, our model incorporates four distinct types of spatiotemporal experts, each specializing in capturing different spatial or temporal dependencies. To reduce the potential computational overhead while integrating multiple experts, we introduce bidirectional spatiotemporal Mamba as experts, each sharing bidirectional temporal and spatial Mamba in distinct combinations to achieve model efficiency and parameter economy. Extensive experiments on four multi-person benchmark datasets demonstrate that our approach not only outperforms state-of-art in accuracy but also reduces model parameter by 41.38% and achieves a 3.6× speedup in training. Zheng Yin, Chengjian Li, Xiangbo Shu, Meiqi Cao, Rui Yan 0010, Jinhui Tang 0001 |
AAAI | 5 |
| 2026 | Spatio-Temporal Decoupled Knowledge Compensator for Few-Shot Action RecognitionabstractFew-Shot Action Recognition (FSAR) is a challenging task that requires recognizing novel action categories with a few labeled videos. Recent works typically apply semantically coarse category names as auxiliary contexts to guide the learning of discriminative visual features. However, such context provided by the action names is too limited to provide sufficient background knowledge for capturing novel spatial and temporal concepts in actions. In this paper, we propose DiST, an innovative Decomposition-incorporation framework for FSAR that makes use of decoupled Spatial and Temporal knowledge provided by large language models to learn expressive multi-granularity prototypes. In the decomposition stage, we decouple vanilla action names into diverse spatio-temporal attribute descriptions (action-related knowledge). Such commonsense knowledge complements semantic contexts from spatial and temporal perspectives. In the incorporation stage, we propose Spatial/Temporal Knowledge Compensators (SKC/TKC) to discover discriminative object-level and frame-level prototypes, respectively. In SKC, object-level prototypes adaptively aggregate important patch tokens under the guidance of spatial knowledge. Moreover, in TKC, frame-level prototypes utilize temporal attributes to assist in inter-frame temporal relation modeling. These learned prototypes thus provide transparency in capturing fine-grained spatial details and diverse temporal patterns. Experimental results show DiST achieves state-of-the-art results on five standard FSAR datasets. Hongyu Qu, Xiangbo Shu, Rui Yan 0010, Hailiang Gao, Wenguan Wang, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Locality-Aware Cross-Modal Correspondence Learning for Dense Audio-Visual Events DetectionabstractDense-localization Audio-Visual Events (DAVE) aims to identify time boundaries and corresponding categories for events that are both audible and visible in a long video, where events may co-occur and exhibit varying durations. However, complex audio-visual scenes often involve asynchronization between modalities, making accurate localization challenging. Existing DAVE solutions extract audio and visual features through unimodal encoders, and fuse them via dense cross-modal interaction. However, independent unimodal encodingstruggles to emphasize shared semantics between modalitieswithout cross-modal guidance, while dense cross-modal attention mayover-attend to semantically unrelated audio-visual features. To address these problems, we present LOCO, a Locality-aware cross-modal Correspondence learning framework for DAVE. LOCO leverages the local temporal continuity of audio-visual events as important guidance to filter irrelevant cross-modal signals and enhance cross-modal alignment throughout both unimodal and cross-modal encoding stages. i) Specifically, LOCO applies Local Correspondence Feature (LCF) Modulation to enforce unimodal encoders to focus on modality-shared semantics by modulating agreement between audio and visual features based on local cross-modal coherence. ii) To better aggregate cross-modal relevant features, we further customize Local Adaptive Cross-modal (LAC) Interaction, which dynamically adjusts attention regions in a data-driven manner. This adaptive mechanism focuses attention on local event boundaries and accommodates varying event durations. By incorporating LCF and LAC, LOCO provides solid performance gains and outperforms existing DAVE methods. The source code will be released. Ling Xing 0003, Hongyu Qu, Rui Yan 0010, Xiangbo Shu, Jinhui Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | VideoExpert: Augmented LLM for Temporal-Sensitive Video UnderstandingabstractThe core challenge in video understanding lies in perceiving dynamic content changes over time. However, multi-modal large language models (MLLMs) struggle with temporal-sensitive video tasks, such as video temporal grounding, which requires generating timestamps to mark the occurrence of specific events. Existing strategies require MLLMs to generate absolute or relative timestamps directly. We have observed that those MLLMs tend to rely more on language patterns than visual cues when generating timestamps, affecting their performance. To address this problem, we propose VideoExpert, a general-purpose MLLM suitable for several temporal-sensitive video tasks. Inspired by the expert concept, VideoExpert integrates two parallel modules: the Temporal Expert and the Spatial Expert. The Temporal Expert is responsible for modeling time sequences and performing temporal grounding. It processes high-frame-rate yet compressed tokens to capture dynamic variations in videos and includes a lightweight prediction head for precise event localization. The Spatial Expert focuses on content detail analysis and instruction following. It handles specially designed spatial tokens and language input, aiming to generate content-related responses. These two experts collaborate seamlessly via a special token, ensuring coordinated temporal grounding and content generation. Notably, the Temporal and Spatial Experts maintain independent parameter sets. This parameter decoupling design enables specialized learning within each part without mutual interference. By offloading temporal grounding from content generation, VideoExpert prevents text pattern biases in timestamp predictions. Moreover, we introduce a Spatial Compress module to obtain spatial tokens. This module filters and compresses patch tokens while preserving key information, delivering compact yet detail-rich input for the Spatial Expert. Extensive experiments conducted on four widely-used benchmarks (i.e. Charades-STA, QVHighlight, YouCookII and NextGQA) across four tasks (temporal grounding, highlight detection, dense video captioning and grounding question answering) demonstrate the effectiveness and versatility of the VideoExpert. Henghao Zhao, Ge-Peng Ji, Rui Yan 0010, Huan Xiong, Zechao Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Attack-Augmented Mixing-Contrastive Skeletal Representation LearningabstractContrastive learning facilitates the acquisition of informative skeleton representations for unsupervised action recognition by leveraging effective positive and negative sample pairs. However, most existing methods construct these pairs through weak or strong data augmentations, which typically rely on random appearance alterations of skeletons. While such augmentations are somewhat effective, they introduce semantic variations only indirectly and face two inherent limitations. First, simply modifying the appearance of skeletons often fails to reflect meaningful semantic variations. Second, random perturbations can unintentionally blur the boundary between positive and negative pairs, weakening the contrastive objective. To address these challenges, we propose an attack-driven augmentation framework that explicitly introduces semantic-level perturbations. This approach facilitates the generation of hard positives while guiding the model to mine more informative hard negatives. Building on this idea, we present Attack-Augmented Mixing-Contrastive Skeletal Representation Learning (A2MC), a novel framework that focuses on contrasting hard positive and hard negative samples for more robust representation learning. Within A2MC, we design an Attack-Augmentation (Att-Aug) module that integrates both targeted (attack-based) and untargeted (augmentation-based) perturbations to generate informative hard positive samples. In parallel, we propose the Positive-Negative Mixer (PNM), which blends hard positive and negative features to synthesize challenging hard negatives. These are then used to update a mixed memory bank for more effective contrastive learning. Comprehensive evaluations across three public benchmarks demonstrate that our approach, termed A2MC, achieves performance on par with or exceeding existing state-of-the-art methods. Binqian Xu, Xiangbo Shu, Jiachao Zhang, Rui Yan 0010, Guosen Xie |
IEEE Trans. Image Process. | 4 |
| 2025 | Exploiting Frequency Dynamics for Enhanced Multimodal Event-Based Action Recognition
Meiqi Cao, Xiangbo Shu, Xin Jiang 0010, Rui Yan 0010, Yazhou Yao, Jinhui Tang 0001 |
ICCV | 4 |
| 2025 | Reliable and Diverse Hierarchical Adapter for Zero-shot Video ClassificationabstractAdapting pre-trained vision-language models to downstream tasks has emerged as a novel paradigm for zero-shot learning. Existing test-time adaptation (TTA) methods such as TPT attempt to fine-tune visual or textual representations to accommodate downstream tasks but still require expensive optimization costs. To this end, Training-free Dynamic Adapter (TDA) maintains a cache containing visual features for each category in a parameter-free manner and measures sample confidence based on prediction entropy of test samples. Inspired by TDA, this work aims to develop the first training-free adapter for zero-shot video classification. Capturing the intrinsic temporal relationships within video data to construct and maintain the video cache is key to extending TDA to the video domain. In this work, we propose a reliable and diverse Hierarchical Adapter for zero-shot video classification, which consists of Frame-level Cache Refiner and Video-level Cache Updater. Before each video sample enters the corresponding cache, it needs to be refined at frame level based on prediction entropy and temporal probability difference. Due to the limited capacity of the cache, we update the cache during inference based on the principle of diversity. Experiments on four popular video classification benchmarks demonstrate the effectiveness of Hierarchical Adapter. The code is available at https://github.com/Gwxer/Hierarchical-Adapter. Wenxuan Ge, Rui Yan 0010, Hongyu Qu, Guosen Xie, Xiangbo Shu |
IJCAI | 3 |
| 2025 | TEST-V: TEst-time Support-set Tuning for Zero-shot Video ClassificationabstractRecently, adapting Vision Language Models (VLMs) to zero-shot visual classification by tuning class embedding with a few prompts (Test-time Prompt Tuning, TPT) or replacing class names with generated visual samples (support-set) has shown promising results. However, TPT cannot avoid the semantic gap between modalities while the support-set cannot be tuned. To this end, we draw on each other's strengths and propose a novel framework, namely TEst-time Support-set Tuning for zero-shot Video Classification (TEST-V). It first dilates the support-set with multiple prompts (Multi-prompting Support-set Dilation, MSD) and then erodes the support-set via learnable weights to mine key cues dynamically (Temporal-aware Support-set Erosion, TSE). Specifically, i) MSD expands the support samples for each class based on multiple prompts inquired from LLMs to enrich the diversity of the support-set. ii) TSE tunes the support-set with factorized learnable weights according to the temporal prediction consistency in a self-supervised manner to dig pivotal supporting cues for each class. TEST-V achieves state-of-the-art results across four benchmarks and shows good interpretability. Rui Yan 0010, Hongyu Qu, Xiaoyu Du 0002, Jinhui Tang 0001, Tieniu Tan |
IJCAI | 1 |
| 2025 | Time-IC: Empowering MLLM with Interleaved Context for Temporal-Sensitive Video UnderstandingabstractTemporal-sensitive video tasks, such as dense video captioning, require models to describe event content in detail and generate timestamps marking their occurrence. Existing approaches tend to over-rely on visuals while neglecting the role of rich contexts, limiting temporal reasoning and semantic comprehension in videos. To this end, we propose Time-IC, a multimodal large language model that incorporates video context for improved understanding. Specifically, Time-IC integrates both internal context (timestamps, speech transcripts) and external context (titles, author descriptions), providing complementary semantic cues and high-level semantic priors. More importantly, it unifies context, video frames, and task instructions into a temporally aligned, interleaved sequence. With instruction tuning, this design enables the model to associate contextual elements and generate comprehensive responses across diverse tasks. Extensive experiments on widely-used benchmarks (YouCookII, QVHighlights) across three tasks (dense video captioning, temporal grounding, highlight detection) demonstrate the effectiveness and flexibility of the proposed Time-IC. Henghao Zhao, Rui Yan 0010, Zechao Li |
MMAsia | 3 |
| 2025 | DAPO : Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage-Based Policy OptimizationabstractThe role of reinforcement learning (RL) in enhancing the reasoning of large language models (LLMs) is becoming increasingly significant. Despite the success of RL in many scenarios, there are still many challenges in improving the reasoning of LLMs. One key challenge is the sparse reward, which introduces more training variance in policy optimization and makes it difficult to obtain a good estimation for value function in Actor-Critic (AC) methods. To address these issues, we introduce Direct Advantage-Based Policy Optimization (DAPO), a novel step-level offline RL algorithm with theoretical guarantees for enhancing the reasoning abilities of LLMs. Unlike response-level methods (such as DPO and GRPO) that the update directions of all reasoning steps are governed by the outcome reward uniformly, DAPO employs a critic function to provide step-level dense signals for policy optimization. Additionally, the actor and critic in DAPO are trained independently, ensuring that critic is a good estimation of true state value function and avoiding the co-training instability observed in standard AC methods. We train DAPO on mathematical and code problems and then evaluate its performance on multiple benchmarks. Our results show that DAPO can effectively enhance the mathematical and code capabilities on both SFT models and RL models, demonstrating the effectiveness of DAPO. Jiacai Liu, Chaojie Wang 0001, Chris Yuhao Liu, Rui Yan 0010, Yang Liu 0084 |
NeurIPS | 5 |
| 2025 | Vision-centric Token Compression in Large Language ModelabstractReal-world applications are stretching context windows to hundreds of thousand of tokens while Large Language Models (LLMs) swell from billions to trillions of parameters.
This dual expansion send compute and memory costs skyrocketing, making $\textit{token compression}$ indispensable.
We introduce Vision Centric Token Compression ($\textbf{Vist}$), a $\textit{slow–fast}$ compression framework that mirrors human reading:
the $\textit{fast}$ path renders distant tokens into images, letting a $\textbf{frozen, lightweight vision encoder}$ skim the low-salience context;
the $\textit{slow}$ path feeds the proximal window into the LLM for fine-grained reasoning.
A Probability-Informed Visual Enhancement (PVE) objective masks high-frequency tokens during training, steering the Resampler to concentrate on semantically rich regions—just as skilled reader gloss over function words.
On eleven in-context learning benchmarks, $\textbf{Vist}$ achieves the same accuracy with 2.3$\times$ fewer tokens, cutting FLOPs by 16\% and memory by 50\%.
This method delivers remarkable results, outperforming the strongest text encoder-based compression method CEPE by $\textbf{7.6}$\% on average over benchmarks like TriviaQA, NQ, PopQA, NLUI, and CLIN, setting a new standard for token efficiency in LLMs. The project is at https://github.com/CSU-JPG/VIST. Ling Xing 0003, Alex Jinpeng Wang, Rui Yan 0010, Xiangbo Shu, Jinhui Tang 0001 |
NeurIPS | 3 |
| 2025 | You Only Communicate Once: One-shot Federated Low-Rank Adaptation of MLLMabstractMultimodal Large Language Models (MLLMs) with Federated Learning (FL) can quickly adapt to privacy-sensitive tasks, but face significant challenges such as high communication costs and increased attack risks, due to their reliance on multi-round communication. To address this, One-shot FL (OFL) has emerged, aiming to complete adaptation in a single client-server communication. However, existing adaptive ensemble OFL methods still need more than one round of communication, because correcting heterogeneity-induced local bias relies on aggregated global supervision, meaning they still do not achieve true one-shot communication. In this work, we make the first attempt to achieve true one-shot communication for MLLMs under OFL, by investigating whether implicit (i.e., initial rather than aggregated) global supervision alone can effectively correct local training bias. Our key finding from the empirical study is that imposing directional supervision on local training substantially mitigates client conflicts and local bias. Building on this insight, we propose YOCO, in which directional supervision with sign-regularized LoRA B enforces global consistency, while sparsely regularized LoRA A preserves client-specific adaptability. Experiments demonstrate that YOCO cuts communication to $\sim$0.03\% of multi-round FL while surpassing those methods in several multimodal scenarios and consistently outperforming all one-shot competitors. Binqian Xu, Haiyang Mei, Zechen Bai, Jinjin Gong, Rui Yan 0010, Guosen Xie, Yazhou Yao, Basura Fernando, Xiangbo Shu |
NeurIPS | 5 |
| 2025 | Test-Time Tuning for Zero-Shot Spatio-Temporal Action Localization
Hongyu Qu, Rui Yan 0010, Jinhui Tang 0001 |
PRCV (11) | 3 |
| 2025 | Appearance-Agnostic Representation Learning for Compositional Action RecognitionabstractThe discussion of compositional generalization in action recognition,i.e., Compositional Action Recognition (CAR), has recently received increasing attention. CAR challenges models to recognize unseen combinations of actions and objects, with the primary challenge being the distribution shift from training to testing. Most previous approaches for CAR incorporate supplementary object annotations (e.g. bounding boxes and objects categories) to learn an instance-centric dynamic representation. However, these methods inevitably introduce stronger visual inductive bias, including object appearance and background bias, that impact generalization performance, particularly in out-of-distribution scenarios. To this end, this work attempts to construct an appearance-agnostic de-biased representation by leveraging the powerful segmentation capability of Segment Anything Model (SAM), which is the first exploration of SAM in the field of compositional action recognition. Specifically, we propose a novel SAM-driven Appearance-Agnostic Representation Learning (A2RL) framework for CAR, which contains two effective sub-modules: Fore-Back Mask (FBM) and Dynamic Relation Modeling (DRM). In FBM, we design a fine-grained instance-invisible and background-removed masking strategy to effectively weaken the strong connection between visual cues and action labels, as well as minimize the impact of irrelevant factors. In DRM, we explore the potential association between subjects and objects involved in one action and then build appearance-agnostic relational descriptors for dynamic modeling. Extensive experiments demonstrate the generalization ability of this work. Notably, FBM achieves significant improvements in all three compositional settings without adding any additional model parameters. The proposed also gains state-of-the-art performance in comparison with the most recent methods in CAR. Xiangbo Shu, Rui Yan 0010, Zhewei Tu, Jinhui Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Eliminating Semantic Ambiguity in Human Pose Estimation via Stable Feature UpsamplingabstractHuman pose estimation is a challenging research task in the computer vision community due to the semantic ambiguity problem caused by inevitable occlusions, varying body shapes, and complex articulations. Although deep learning-based methods have significantly improved the performance of this task, existing feature upsampling operations,e.g., bilinear interpolation and transposed convolution, within current convolutional neural networks and Transformer frameworks suffer from a multitude of limitations, including the inability to adapt to specific tasks and the loss of fine-grained semantic details. In this work, we propose a simple yet effective two-step stable feature upsampling (SIU) strategy that addresses these limitations by leveraging a learnable and efficient upsampling operation. Specifically, wefirstapply periodic shuffling to increase the resolution of the feature maps.Secondly, we utilize convolution layers to adjust the size of feature channels to match those of the input feature maps. The proposed SIU enables the entire network to adapt to the specific feature requirements of the human pose estimation task, making it more effective in preserving spatial information. Quantitatively, extensive experimental results on the challenging COCO-WholeBody dataset validate that our approach outperforms state-of-the-art methods accurately and efficiently, and possesses strong transferability, making it applicable to a wide range of baselines. Moreover, the qualitative results validate that SIU can effectively eliminate the semantic ambiguity problem in challenging pose scenarios, such as occlusions and overlapping1. Rui Yan 0010, Xiangbo Shu, Pingcheng Dong, Long Chen 0016, Xiaoyu Du 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | STPM: Spatial-Temporal Token Pruning and Merging for Complex Activity RecognitionabstractLightweight video representation techniques have advanced significantly for simple activity recognition, but they still encounter several issues when applied to complex activity recognition: 1) The presence of numerous individuals and varying spatial positions makes it difficult for traditional token pruning methods to maintain accuracy. 2) Simply discarding entire frames may result in the loss of crucial clues. 3) To maintain parallel computing, applying the same pruning rate to every frame leads to significant redundancy in frames with low information content. To this end, we propose a lightweight and novel Spatial-Temporal Token Pruning and Merging (STPM) framework, specifically designed for complex action videos where human actors occupy a small spatial resolution within video frames. Our framework considers two critical factors: semantic importance and spatial-temporal redundancy, to further reduce overhead. For semantic importance, STPM captures class-specific attention scores by learning multiple class tokens within the transformer to guide token pruning. For spatial-temporal redundancy, STPM employs an anchor graph and temporal attention to perform spatial and temporal token merging, preserving appearance and temporal cues while eliminating semantic duplication and redundancy. We conduct extensive experiments on JRDB-PAR primarily using recently introduced video transformer backbones, e.g., MViT and ViT. Our framework achieves similar results while requiring 40% less computation. Yumeng Su, Jiachao Zhang, Rui Yan 0010, Pengpeng Li 0001, Guosen Xie, Xiangbo Shu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | GPT4Ego: Unleashing the Potential of Pre-Trained Models for Zero-Shot Egocentric Action RecognitionabstractVision-Language Models (VLMs), pre-trained on large-scale datasets, have shown impressive performance in various visual recognition tasks. This advancement paves the way for notable performance in some egocentric tasks, Zero-Shot Egocentric Action Recognition (ZS-EAR), entailing VLMs zero-shot to recognize actions from first-person videos enriched in more realistic human-environment interactions. Typically, VLMs handle ZS-EAR as a global video-text matching task, which often leads to suboptimal alignment of vision and linguistic knowledge. We propose a refined approach for ZS-EAR using VLMs, emphasizing fine-grained concept-description alignment that capitalizes on the rich semantic and contextual details in egocentric videos. In this work, we introduce a straightforward yet remarkably potent VLM framework,akaGPT4Ego, designed to enhance the fine-grained alignment of concept and description between vision and language. Specifically, we first propose a new Ego-oriented Text Prompting (EgoTP$\spadesuit$) scheme, which effectively prompts action-related text-contextual semantics by evolving word-level class names to sentence-level contextual descriptions by ChatGPT with well-designed chain-of-thought textual prompts. Moreover, we design a new Ego-oriented Visual Parsing (EgoVP$\clubsuit$) strategy that learns action-related vision-contextual semantics by refining global-level images to part-level contextual concepts with the help of SAM. Extensive experiments demonstrate GPT4Ego significantly outperforms existing VLMs on three large-scale egocentric video benchmarks, i.e., EPIC-KITCHENS-100 (33.2%$\uparrow$$_{\bm {+9.4}}$), EGTEA (39.6%$\uparrow$$_{\bm {+5.5}}$), and CharadesEgo (31.5%$\uparrow$$_{\bm {+2.6}}$). In addition, benefiting from the novel mechanism of fine-grained concept and description alignment, GPT4Ego can sustainably evolve with the advancement of ever-growing pre-trained foundational models. We hope this work can encourage the egocentric community to build more investigation into pre-trained vision-language models. Guangzhao Dai, Xiangbo Shu, Rui Yan 0010, Jiachao Zhang |
IEEE Trans. Multim. | 4 |
| 2025 | Hierarchical Motion-Enhanced Matching Framework for Few-Shot Action RecognitionabstractFew-Shot Action Recognition (FSAR) aims to recognize novel class action with limited annotated training data from the same class. Most FSAR methods subconsciously follow the few-shot image classification solutions by solely focusing on appearance-level matching between support and query videos, such as part-level matching, frame-level matching, and segment-level matching. However, these methods, almost always, have two main limitations: 1) generally ignore the relationship among these part-, frame- and segment-level features and 2) may mismatch the same class actions under fast-term and slow-term dynamics. To this end, we present a novel Hierarchical Motion-enhanced Matching (HM${^{2}}$) framework to hierarchically learn the relation-aware multi-modal features, and jointly promote the multi-modal matching, including appearance-level matching on segments, frames, and parts, as well as the motion-level matching on dynamics. Specifically, we first propose a new Hierarchical Tokenizer (HT) to learn multi-modal features, namely utilizing a hierarchical Transformer to learn appearance-level features, along with a Slow-Fast Aware Motion (SFAM) strategy to learn motion-level features covering fast- and slow-term dynamics. Next, we propose a new Relation-aware Matcher (RM) to match the multi-modal features, by leveraging a Hierarchical Relational Graph Convolutional Network (H-RGCN) to capture the relationship among these appearance-level features. Further, a Dual Sample-to-Class Matching (DSCM) strategy is proposed to measure the bidirectional similarities among appearance- and motion-modal features by sample-to-class matching and class-to-sample matching. Extensive experiments on four golden FSAR datasets demonstrate significant performance improvements of HM${^{2}}$compared with the state-of-the-art methods. Hailiang Gao, Guosen Xie, Rui Yan 0010, Qiongjie Cui, Hongyu Qu, Xiangbo Shu |
IEEE Trans. Multim. | 3 |
| 2025 | MVP-Shot: Multi-Velocity Progressive-Alignment Framework for Few-Shot Action RecognitionabstractRecent few-shot action recognition (FSAR) methods typically perform semantic matching on learned discriminative features to achieve promising performance. However, most FSAR methods focus on single-scale (e.g., frame-level, segment-level,etc.) feature alignment, which ignores that human actions with the same semantic may appear at different velocities. To this end, we develop a novel Multi-Velocity Progressive-alignment (MVP-Shot) framework to progressively learn and align semantic-related action features at multi-velocity levels. Concretely, a Multi-Velocity Feature Alignment (MVFA) module is designed to measure the similarity between features from support and query videos with different velocity scales and then merge all similarity scores in a residual fashion. To avoid the multiple velocity features deviating from the underlying motion semantic, our proposed Progressive Semantic-Tailored Interaction (PSTI) module injects velocity-tailored text information into the video feature via feature interaction on channel and temporal domains at different velocities. The above two modules compensate for each other to make more accurate query sample predictions under the few-shot settings. Experimental results show our method outperforms current state-of-the-art methods on multiple standard few-shot benchmarks (i.e., HMDB51, UCF101, Kinetics, SSv2-full, and SSv2-small). Hongyu Qu, Rui Yan 0010, Xiangbo Shu, Hailiang Gao, Guosen Xie |
IEEE Trans. Multim. | 2 |
| 2025 | Coarse-Fine Nested Network for Weakly Supervised Group Activity RecognitionabstractWeakly supervised group activity recognition (WSGAR) aims at identifying the overall behavior of multiple persons without any fine-grained supervision information (including individual position and action label). Traditional methods usually adopt a person-to-whole way: detect persons via off-the-shelf detectors, obtain person-level features, and integrate into the group-level features for training the classifier. However, these methods are unflexible due to serious reliance on the quality of detectors. To get rid of the detector, recent works learn several prototype tokens from noisy grid features with learnable weights directly, which treat all the local visual information equally and bring in redundant and ambiguous information to some extent. To this end, we propose a novel coarse-fine nested network (CFNN) to coarsely localize the key visual patches of activity and further finely learn the local features, as well as the global features. Specifically, we design a nested interactor (NI) to progressively model the spatiotemporal interactions of the learnable global token. According to the cue of spatial interaction in NI, we localize several key visual patches via a new coarse-grained spatial localizer (CSL). Then, we finally encode these localized visual patches with the help of global spatiotemporal dependency via a new fine-grained spatiotemporal selector (FSS). Extensive experiments on Volleyball and NBA datasets demonstrate the effectiveness of the proposed CFNN compared with the existing competitive methods. Code is available at: https://github.com/gexiaojingshelby/CFNN. Xiaojing Ge, Rui Yan 0010, Xiangbo Shu, Keke Chen, Guosen Xie |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Leveraging Frame- and Feature-level Progressive Augmentation for Semi-supervised Action RecognitionabstractSemi-supervised action recognition is a challenging yet prospective task due to its low reliance on costly labeled videos. One high-profile solution is to explore frame-level weak/strong augmentations for learning abundant representations, inspired by the FixMatch framework dominating the semi-supervised image classification task. However, such a solution mainly brings perturbations in terms of texture and scale, leading to the limitation in learning action representations in videos with spatiotemporal redundancy and complexity. Therefore, we revisit the creative trick of weak/strong augmentations in FixMatch and then propose, to the best of our knowledge, a novel Frame- and Feature-level augmentation FixMatch (dubbed as F 2 -FixMatch) framework to learn more abundant action representations for being robust to complex and dynamic video scenarios. Specifically, we design a new Progressive Augmentation mechanism that implements the weak/strong augmentations first at the frame level, and further implements the perturbation at the feature level, to obtain abundant four types of augmented features in broader perturbation spaces. Moreover, we present an evolved Multihead Pseudo-Labeling scheme to promote the consistency of features across different augmented versions based on the pseudo labels. We conduct extensive experiments on several public datasets to demonstrate that our F 2 -FixMatch achieves the performance gain compared with current state-of-the-art methods. The source codes of F 2 -FixMatch are publicly available at https://github.com/zwtu/F2FixMatch . Zhewei Tu, Xiangbo Shu, Rui Yan 0010, Zhenxing Liu 0001, Jiachao Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | DTS-TPT: Dual Temporal-Sync Test-time Prompt Tuning for Zero-shot Activity Recognition
Rui Yan 0010, Hongyu Qu, Xiangbo Shu, Jinhui Tang 0001, Tieniu Tan |
IJCAI | 1 |
| 2024 | AdaFPP: Adapt-Focused Bi-Propagating Prototype Learning for Panoramic Activity RecognitionabstractPanoramic Activity Recognition (PAR) aims to identify multi-granul-arity behaviors performed by multiple persons in panoramic scenes, including individual activities, group activities, and global activities. Previous methods 1) heavily rely on manually annotated detection boxes in training and inference, hindering further practical deployment; or 2) directly employ normal detectors to detect multiple persons with varying size and spatial occlusion in panoramic scenes, blocking the performance gain of PAR. To this end, we consider learning a detector adapting varying-size occluded persons, which is optimized along with the recognition module in the all-in-one framework. Therefore, we propose a novel Adapt-Focused bi-Propagating Prototype learning (AdaFPP) framework to jointly recognize individual, group, and global activities in panoramic activity scenes by learning an adapt-focused detector and multi-granularity prototypes as the pretext tasks in an end-to-end way. Specifically, to accommodate the varying sizes and spatial occlusion of multiple persons in crowed panoramic scenes, we introduce a panoramic adapt-focuser, achieving the size-adapting detection of individuals by comprehensively selecting and performing fine-grained detections on object-dense sub-regions identified through original detections. In addition, to mitigate information loss due to inaccurate individual localizations, we introduce a bi-propagation prototyper that promotes closed-loop interaction and informative consistency across different granularities by facilitating bidirectional information propagation among the individual, group, and global levels. Extensive experiments demonstrate the significant performance of AdaFPP and emphasize its powerful applicability for PAR. Meiqi Cao, Rui Yan 0010, Xiangbo Shu, Guangzhao Dai, Yazhou Yao, Guosen Xie |
ACM Multimedia | 2 |
| 2024 | DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval GuidelinesabstractFine-grained image retrieval (FGIR) is to learn visual representations that distinguish visually similar objects while maintaining generalization. Existing methods propose to generate discriminative features, but rarely consider the particularity of the FGIR task itself. This paper presents a meticulous analysis leading to the proposal of practical guidelines to identify subcategory-specific discrepancies and generate discriminative features to design effective FGIR models. These guidelines include emphasizing the object (G1), highlighting subcategory-specific discrepancies (G2), and employing effective training strategy (G3). Following G1 and G2, we design a novel Dual Visual Filtering mechanism for the plain visual transformer, denoted as DVF, to capture subcategory-specific discrepancies. Specifically, the dual visual filtering mechanism comprises an object-oriented module and a semantic-oriented module. These components serve to magnify objects and identify discriminative regions, respectively. Following G3, we implement a discriminative model training strategy to improve the discriminability and generalization ability of DVF. Extensive analysis and ablation studies confirm the efficacy of our proposed guidelines. Without bells and whistles, the proposed DVF achieves state-of-the-art performance on three widely-used fine-grained datasets in closed-set and open-set settings. Xin Jiang 0010, Hao Tang 0007, Rui Yan 0010, Jinhui Tang 0001, Zechao Li |
ACM Multimedia | 3 |
| 2024 | Boosting Weakly-Supervised Image Segmentation via Representation, Transform, and CompensatorabstractWeakly-supervised image segmentation (WSIS) is a fundamental task in the domain of computer vision that relies on image-level class labels. While multi-stage training procedures have been widely used in existing WSIS methods to obtain high-quality pseudo-masks as ground-truth, resulting in significant progress, single-stage WSIS methods have recently gained attention due to their potential for simplifying the training procedure. However, single-stage methods suffer from low-quality pseudo-masks that limit their practical applications. To address this problem, this paper proposes a novel single-stage WSIS method that utilizes a siamese network with contrastive learning to improve the quality of class activation maps (CAMs) and achieve a self-refinement result. The proposed method employs a cross-representation refinement method that expands reliable object regions by utilizing different feature representations from the backbone. Besides, a cross-transform regularization module is introduced that learns robust class prototypes for contrastive learning and captures global context information to feed back rough CAMs, thereby improving the quality of CAMs. The final high-quality CAMs are used as pseudo-masks to supervise the segmentation result. Experimental results on the PASCAL VOC 2012 and COCO datasets demonstrate that the proposed method significantly outperforms other state-of-the-art methods, achieving 72.38% and 72.95% mIoU on PASCAL VOC 2012 val set and test set, 42.51% mIoU on COCO val set, respectively. Furthermore, the proposed method has been extended to weakly supervised object localization, and experimental results demonstrate that it continues to achieve very competitive results. The source codes have been released athttps://github.com/ChunyanWang1/RTC. Rui Yan 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Semantic-Disentangled Transformer With Noun-Verb Embedding for Compositional Action RecognitionabstractRecognizing actions performed on unseen objects, known as Compositional Action Recognition (CAR), has attracted increasing attention in recent years. The main challenge is to overcome the distribution shift of "action-objects" pairs between the training and testing sets. Previous works for CAR usually introduce extra information (e.g. bounding box) to enhance the dynamic cues of video features. However, these approaches do not essentially eliminate the inherent inductive bias in the video, which can be regarded as the stumbling block for model generalization. Because the video features are usually extracted from the visually cluttered areas in which many objects cannot be removed or masked explicitly. To this end, this work attempts to implicitly accomplish semantic-level decoupling of "object-action" in the high-level feature space. Specifically, we propose a novel Semantic-Decoupling Transformer framework, dubbed as DeFormer, which contains two insightful sub-modules: Objects-Motion Decoupler (OMD) and Semantic-Decoupling Constrainer (SDC). In OMD, we initialize several learnable tokens incorporating annotation priors to learn an instance-level representation and then decouple it into the appearance feature and motion feature in high-level visual space. In SDC, we use textual information in the high-level language space to construct a dual-contrastive association to constrain the decoupled appearance feature and motion feature obtained in OMD. Extensive experiments verify the generalization ability of DeFormer. Specifically, compared to the baseline method, DeFormer achieves absolute improvements of 3%, 3.3%, and 5.4% under three different settings on STH-ELSE, while corresponding improvements on EPIC-KITCHENS-55 are 4.7%, 9.2%, and 4.4%. Besides, DeFormer gains state-of-the-art results either on ground-truth or detected annotations. Rui Yan 0010, Xiangbo Shu, Zhewei Tu, Guangzhao Dai, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Video-Text Pre-training with Learned Regions for RetrievalabstractVideo-Text pre-training aims at learning transferable representations from large-scale video-text pairs via aligning the semantics between visual and textual information. State-of-the-art approaches extract visual features from raw pixels in an end-to-end fashion. However, these methods operate at frame-level directly and thus overlook the spatio-temporal structure of objects in video, which yet has a strong synergy with nouns in textual descriptions. In this work, we propose a simple yet effective module for video-text representation learning, namely RegionLearner, which can take into account the structure of objects during pre-training on large-scale video-text pairs. Given a video, our module (1) first quantizes continuous visual features via clustering patch-features into the same cluster according to content similarity, then (2) generates learnable masks to aggregate fragmentary features into regions with complete semantics, and finally (3) models the spatio-temporal dependencies between different semantic regions. In contrast to using off-the-shelf object detectors, our proposed module does not require explicit supervision and is much more computationally efficient. We pre-train the proposed approach on the public WebVid2M and CC3M datasets. Extensive evaluations on four downstream video-text retrieval benchmarks clearly demonstrate the effectiveness of our RegionLearner. Rui Yan 0010, Zheng Shou 0001, Yixiao Ge, Jinpeng Wang 0001, Xudong Lin 0003, Guanyu Cai, Jinhui Tang 0001 |
AAAI | 1 |
| 2023 | M3Net: Multi-view Encoding, Matching, and Fusion for Few-shot Fine-grained Action RecognitionabstractDue to the scarcity of manually annotated data required for fine-grained video understanding, few-shot fine-grained (FS-FG) action recognition has gained significant attention, with the aim of classifying novel fine-grained action categories with only a few labeled instances. Despite the progress made in FS coarse-grained action recognition, current approaches encounter two challenges when dealing with the fine-grained action categories: the inability to capture subtle action details and the insufficiency of learning from limited data that exhibit high intra-class variance and inter-class similarity. To address these limitations, we propose M3Net, a matching-based framework for FS-FG action recognition, which incorporates multi-view encoding, multi-view matching, and multi-view fusion to facilitate embedding encoding, similarity matching, and decision making across multiple viewpoints.Multi-view encoding captures rich contextual details from the intra-frame, intra-video, and intra-episode perspectives, generating customized higher-order embeddings for fine-grained data.Multi-view matching integrates various matching functions enabling flexible relation modeling within limited samples to handle multi-scale spatio-temporal variations by leveraging the instance-specific, category-specific, and task-specific perspectives. Multi-view fusion consists of matching-predictions fusion and matching-losses fusion over the above views, where the former promotes mutual complementarity and the latter enhances embedding generalizability by employing multi-task collaborative learning. Explainable visualizations and experimental results on three challenging benchmarks demonstrate the superiority of M3Net in capturing fine-grained action details and achieving state-of-the-art performance for FS-FG action recognition. Hao Tang 0007, Jun Liu 0036, Shuanglin Yan, Rui Yan 0010, Zechao Li, Jinhui Tang 0001 |
ACM Multimedia | 4 |
| 2023 | MUP: Multi-granularity Unified Perception for Panoramic Activity RecognitionabstractPanoramic activity recognition is required to jointly identify multi-granularity human behaviors including individual actions, group activities, and global activities in multi-person videos. Previous methods encode these behaviors hierarchically through multiple stages, which disturb the inherent co-occurrence across multi-granularity behaviors in the same scene. To this end, we propose a novel Multi-granularity Unified Perception (MUP) framework that perceives different granularity behaviors universally to explore the co-occurrence motion pattern via the same parameters in an end-to-end fashion. To be specific, the proposed framework stacks three Unified Motion Encoding (UME) blocks for modeling multiple granularity behaviors with shared parameters. UME block mines intra-relevant and cross-relevant semantics synchronously from input feature sequences via Intra-granularity Motion Embedding (IME) and Cross-granularity Motion Prototyping (CMP). In particular, IME aims to model the interactions among visual features within each granularity based on the attention mechanism. CMP aims to aggregate features across different granularities (i.e., person to group) via several learnable prototypes. Extensive experiments demonstrate that MUP outperforms the state-of-the-art methods on JRDB-PAR and has satisfactory interpretability. Meiqi Cao, Rui Yan 0010, Xiangbo Shu, Jiachao Zhang, Jinpeng Wang 0001, Guosen Xie |
ACM Multimedia | 2 |
| 2023 | Foreground/Background-Masked Interaction Learning for Spatio-temporal Action DetectionabstractSpatio-temporal Action Detection (SAD) aims to recognize the multi-class actions, and meanwhile locate their spatio-temporal occurrence in untrimmed videos. Besides relying on the inherent inter-actor interactions, most previous SAD approaches model actor interactions between multi-actors and the whole frames or special parts (e.g., objects/hands). However, such approaches are relatively graceless by 1) roughly treating all various actors to equivalently interact with frames/parts or by 2) sumptuously borrowing multiple costly detectors to acquire the special parts. To solve the above dilemma, we propose a novel Foreground/Background-masked Interaction Learning (dubbed as FBI Learning) framework to learn the multi-actor features by attentively interacting with the hands-down foreground and background frames. Specifically, we first design a new Mask-guided Cross Attention (MCA) mechanism that calculates the masked cross-attentions to capture the compact relations between the actors and foreground/background regions. Next, we present a new Actor-guided Feature Aggregation (AFA) scheme that integrates foreground- and background-interacted actor features with the learnable actor-based weights. Finally, we construct a long-term feature bank that associates temporal context information to facilitate action classification. Extensive experiments are conducted on commonly available UCF101-24, MultiSports, and AVA v2.1/v2.2 datasets, which illustrate the competitive performance of FBI Learning against the state-of-the-art methods. Keke Chen, Xiangbo Shu, Guosen Xie, Rui Yan 0010, Jinhui Tang 0001 |
ACM Multimedia | 4 |
| 2023 | Slowfast Diversity-aware Prototype Learning for Egocentric Action RecognitionabstractEgocentric Action Recognition (EAR) is required to recognize both the interacting objects (noun) and the motion (verb) against cluttered backgrounds with distracting objects. For capturing interacting objects, traditional approaches heavily rely on luxury object annotations or detectors, though a few works heuristically enumerate the fixed sets of verb-constrained prototypes to roughly exclude the background. For capturing motion, the inherent variations of motion duration among egocentric videos with different lengths are almost ignored. To this end, we propose a novel Slowfast Diversity-aware Prototype learning (SDP) to effectively capture interacting objects by learning compact yet diverse prototypes, and adaptively capture motion in either long-time video or short-time video. Specifically, we present a new Part-to-Prototype (P2P) scheme to learn prototypes from raw videos covering the interacting objects by refining the semantic information from part level to prototype level. Moreover, for adaptively capturing motion, we design a new Slow-Fast Context (SFC) mechanism that explores the Up/Down augmentations for the prototype representation at the semantic level to strengthen the transient dynamic information in short-time videos and eliminate the redundant dynamic information in long-time videos, which are further fine-complemented via the slow- and fast-aware attentions. Extensive experiments demonstrate SDP outperforms state-of-the-art methods on two large-scale egocentric video benchmarks, i.e., EPIC-KITCHENS-100 and EGTEA. Guangzhao Dai, Xiangbo Shu, Rui Yan 0010, Jinhui Tang 0001 |
ACM Multimedia | 3 |
| 2023 | Progressive Instance-Aware Feature Learning for Compositional Action RecognitionabstractIn order to enable the model to generalize to unseen "action-objects" (compositional action), previous methods encode multiple pieces of information (i.e., the appearance, position, and identity of visual instances) independently and concatenate them for classification. However, these methods ignore the potential supervisory role of instance information (i.e., position and identity) in the process of visual perception. To this end, we present a novel framework, namely Progressive Instance-aware Feature Learning (PIFL), to progressively extract, reason, and predict dynamic cues of moving instances from videos for compositional action recognition. Specifically, this framework extracts features from foreground instances that are likely to be relevant to human actions (Position-aware Appearance Feature Extraction in Section III-B1), performs identity-aware reasoning among instance-centric features with semantic-specific interactions (Identity-aware Feature Interaction in Section III-B2), and finally predicts instances' position from observed states to force the model into perceiving their movement (Semantic-aware Position Prediction in Section III-B3). We evaluate our approach on two compositional action recognition benchmarks, namely, Something-Else and IKEA-Assembly. Our approach achieves consistent accuracy gain beyond off-the-shelf action recognition algorithms in terms of both ground truth and detected position of instances. Rui Yan 0010, Lingxi Xie, Xiangbo Shu, Liyan Zhang 0001, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | HiGCIN: Hierarchical Graph-Based Cross Inference Network for Group Activity RecognitionabstractGroup activity recognition (GAR) is a challenging task aimed at recognizing the behavior of a group of people. It is a complex inference process in which visual cues collected from individuals are integrated into the final prediction, being aware of the interaction between them. This paper goes one step further beyond the existing approaches by designing a Hierarchical Graph-based Cross Inference Network (HiGCIN), in which three levels of information, i.e., the body-region level, person level, and group-activity level, are constructed, learned, and inferred in an end-to-end manner. Primarily, we present a generic Cross Inference Block (CIB), which is able to concurrently capture the latent spatiotemporal dependencies among body regions and persons. Based on the CIB, two modules are designed to extract and refine features for group activities at each level. Experiments on two popular benchmarks verify the effectiveness of our approach, particularly in the ability to infer with multilevel visual cues. In addition, training our approach does not require individual action labels to be provided, which greatly reduces the amount of labor required in data annotation. Rui Yan 0010, Lingxi Xie, Jinhui Tang 0001, Xiangbo Shu, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Com-STAL: Compositional Spatio-Temporal Action LocalizationabstractSpatio-temporal action localization aims to locate the spatial and temporal positions of actors and classify their actions. However, prior research overlooks the fact that human actions often interact with novel objects in real-world scenarios, which neglects the various combinations of action-object, and considerably limits the generalization of the developed models. In this paper, we study the action-object combinations by researching multi-modal vision information of them. To this end, we propose a novel compositional spatio-temporal action localization (Com-STAL) task, which features non-overlapping action-object combinations in their training and test sets. Based on this, we construct a compositional action localization dataset (Com-AD). Beyond that, we propose a simple yet effective framework, Instance-Centric Interaction Network (ICIN), to reduce invalid induction biases within the visual modality and alleviate the combined distribution bias issue by leveraging additional modal information. The extensive experiment results on Com-AD demonstrate superior action localization performance of ICIN. Shaomeng Wang, Rui Yan 0010, Guangzhao Dai, Yan Song 0005, Xiangbo Shu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Magi-Net: Meta Negative Network for Early Activity PredictionabstractEarly activity prediction/recognition aims to recognize action categories before they are fully conveyed. Compared to full-length action sequences, partial video sequences only provide insufficient discrimination information, which makes predicting the class labels for some similar activities challenging, especially when only very few frames can be observed. To address this challenge, in this paper, we propose a novel meta negative network, namely, Magi-Net, that utilizes a contrastive learning scheme to alleviate the insufficiency of discriminative information. In our Magi-Net model, the positive samples are generated by augmenting an input anchor conditioned on all observation ratios, while the negative samples are selected from a trainable negative look-up memory (LUM) table, which stores the training samples and the corresponding misleading categories. Furthermore, a meta negative sample optimization strategy (MetaSOS) is proposed to boost the training of Magi-Net by encouraging the model to learn from the most informative negative samples via a meta learning scheme. Extensive experiments are conducted on several public skeleton-based activity datasets, and the results show the efficacy of the proposed Magi-Net model. Faliang Chang, Junhao Zhang 0001, Rui Yan 0010, Chunsheng Liu 0001, Bin Wang 0004, Zheng Shou 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Look Less Think More: Rethinking Compositional Action RecognitionabstractCompositional action recognition which aims to identify the unseen combinations of actions and objects has recently attracted wide attention. Conventional methods bring in additional cues (e.g., dynamic motions of objects) to alleviate the inductive bias between the visual appearance of objects and the human action-level labels. Besides, compared with non-compositional settings, previous methods only pursue higher performance in compositional settings, which can not prove their generalization ability. To this end, we firstly rethink the problem and design a more generalized metric (namely Drop Ratio) and a more practical setting to evaluate the compositional generalization of existing action recognition algorithms. Beyond that, we propose a simple yet effective framework, Look Less Think More (LLTM), to reduce the strong association between visual objects and action-level labels (Look Less), and then discover the commonsense relationships between object categories and human actions (Think More). We test the rationality of the proposed Drop Ratio and Practical setting by comparing several popular action recognition methods on SSV2. Besides, the proposed LLTM achieves state-of-the-art performance on SSV2 with different settings. Rui Yan 0010, Xiangbo Shu, Junhao Zhang 0001, Yonghua Pan, Jinhui Tang 0001 |
ACM Multimedia | 1 |
| 2022 | Spatiotemporal Perturbation Based Dynamic Consistency for Semi-supervised Temporal Action Detection
Yan Song 0005, Rui Yan 0010, Xiangbo Shu |
MMM (1) | 3 |
| 2022 | Skip-attention encoder-decoder framework for human motion prediction
Xiangbo Shu, Rui Yan 0010, Jiachao Zhang, Yan Song 0005 |
Multim. Syst. | 3 |
| 2022 | Coherence Constrained Graph LSTM for Group Activity RecognitionabstractThis work aims to address the group activity recognition problem by exploring human motion characteristics. Traditional methods hold that the motions of all persons contribute equally to the group activity, which suppresses the contributions of some relevant motions to the whole activity while overstating some irrelevant motions. To address this problem, we present a Spatio-Temporal Context Coherence (STCC) constraint and a Global Context Coherence (GCC) constraint to capture the relevant motions and quantify their contributions to the group activity, respectively. Based on this, we propose a novel Coherence Constrained Graph LSTM (CCG-LSTM) with STCC and GCC to effectively recognize group activity, by modeling the relevant motions of individuals while suppressing the irrelevant motions. Specifically, to capture the relevant motions, we build the CCG-LSTM with a temporal confidence gate and a spatial confidence gate to control the memory state updating in terms of the temporally previous state and the spatially neighboring states, respectively. In addition, an attention mechanism is employed to quantify the contribution of a certain motion by measuring the consistency between itself and the whole activity at each time step. Finally, we conduct experiments on two widely-used datasets to illustrate the effectiveness of the proposed CCG-LSTM compared with the state-of-the-art methods. Jinhui Tang 0001, Xiangbo Shu, Rui Yan 0010, Liyan Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Expansion-Squeeze-Excitation Fusion Network for Elderly Activity RecognitionabstractThis work focuses on the task of elderly activity recognition, which is a challenging task due to the existence of individual actions and human-object interactions in elderly activities. Thus, we attempt to effectively aggregate the discriminative information of actions and interactions from both RGB videos and skeleton sequences by attentively fusing multi-modal features. Recently, some nonlinear multi-modal fusion approaches are proposed by utilizing nonlinear attention mechanism that is extended from Squeeze-and-Excitation Networks (SENet). Inspired by this, we propose a novel Expansion-Squeeze-Excitation Fusion Network (ESE-FN) to effectively address the problem of elderly activity recognition, which learns modal and channel-wise Expansion-Squeeze-Excitation (ESE) attentions for attentively fusing the multi-modal features in the modal and channel-wise ways. Specifically, ESE-FN firstly implements the modal-wise fusion with the Modal-wise ESE Attention (M-ESEA) to aggregate discriminative information in modal-wise way, and then implements the channel-wise fusion with the Channel-wise ESE Attention (C-ESEA) to aggregate the multi-channel discriminative information in channel-wise way (referring toFigure 1). Furthermore, we design a new Multi-modal Loss (ML) to keep the consistency between the single-modal features and the fused multi-modal features by adding the penalty of difference between the minimum prediction losses on single modalities and the prediction loss on the fused modality. Finally, we conduct experiments on a largest-scale elderly activity dataset, i.e., ETRI-Activity3D (including 110,000+ videos, and 50+ categories), to demonstrate that the proposed ESE-FN achieves the best accuracy compared with the state-of-the-art methods. In addition, more extensive experimental results show that the proposed ESE-FN is also comparable to the other methods in terms of normal action recognition task. Xiangbo Shu, Jiawen Yang, Rui Yan 0010, Yan Song 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Position-Aware Participation-Contributed Temporal Dynamic Model for Group Activity RecognitionabstractGroup activity recognition (GAR) aiming at understanding the behavior of a group of people in a video clip has received increasing attention recently. Nevertheless, most of the existing solutions ignore that not all the persons contribute to the group activity of the scene equally. That is to say, the contribution from different individual behaviors to group activity is different; meanwhile, the contribution from people with different spatial positions is also different. To this end, we propose a novel Position-aware Participation-Contributed Temporal Dynamic Model (P2CTDM), in which two types of the key actor are constructed and learned. Specifically, we focus on the behaviors of key actors, who maintain steady motions (long moving time, called long motions) or display remarkable motions (but closely related to other people and the group activity, called flash motions) at a certain moment. For capturing long motions, we rank individual motions according to their intensity measured by stacking optical flows. For capturing flash motions that are closely related to other people, we design a position-aware interaction module (PIM) that simultaneously considers the feature similarity and position information. Beyond that, for capturing flash motions that are highly related to the group activity, we also present an aggregation long short-term memory (Agg-LSTM) to fuse the outputs from PIM by time-varying trainable attention factors. Four widely used benchmarks are adopted to evaluate the performance of the proposed P2CTDM compared to the state of the art. Rui Yan 0010, Xiangbo Shu, Chengcheng Yuan, Qi Tian 0001, Jinhui Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Social Adaptive Module for Weakly-Supervised Group Activity Recognition
Rui Yan 0010, Lingxi Xie, Jinhui Tang 0001, Xiangbo Shu, Qi Tian 0001 |
ECCV (8) | 1 |
| 2020 | Storyboard relational model for group activity recognitionabstractThis work concerns how to effectively recognize the group activity performed by multiple persons collectively. As known, Storyboards (i.e., medium shot, close shot) jointly describe the whole storyline of a movie in a compact way. Likewise, the actors in small subgroups (similar to Storyboards) of a group activity scene contribute a lot to such group activity and develop more compact relationships among them within subgroups. Inspired by this, we propose a Storyboard Relational Model (SRM) to address the problem of Group Activity Recognition by splitting and reintegrating the group activity based on the small yet compact Storyboards. SRM mainly consists of a Pose-Guided Pruning (PGP) module and a Dual Graph Convolutional Networks (Dual-GCN) module. Specifically, PGP is designed to refine a series of Storyboards from the group activity scene by leveraging the attention ranges of individuals. Dual-GCN models the compact relationships among actors in a Storyboard. Experimental results on two widely-used datasets illustrate the effectiveness of the proposed SRM compared with the state-of-the-art methods. Boning Li, Xiangbo Shu, Rui Yan 0010 |
MMAsia | 3 |
| 2018 | Participation-Contributed Temporal Dynamic Model for Group Activity RecognitionabstractGroup activity recognition, a challenging task that a number of individuals occur in the scene of activity while only a small subset of them participate in, has received increasing attentions. However, most of the previous methods model all the individuals' actions equivalently while ignoring a fact that not all of them are contributed to the discrimination of group activity. That is to say, only a small number of key actors (participants) play important roles in the whole group activity. Inspired by this, we explore a new "One to Key" idea to progressively aggregate temporal dynamics of key actors with different participation degrees over time from each person. Here, we focus on two types of key actors in the whole activity, who steadily move in the whole process (long moving time) or intensely move (but closely related to the group activity) at a significant moment. Based on this, we propose a novel Participation-Contributed Temporal Dynamic Model (PC-TDM) to recognize group activity, which mainly consists of a "One" network and a "One to Key" network. Specifically, "One" network aims at modeling the individual dynamic of each person. "One to Key" network feeds the outputs from the "One" network into a Bidirectional LSTM (Bi-LSTM) according to the order of individual's moving time. Subsequently, each output state of Bi-LSTM weighted by a trainable time-varying attention factor is aggregated by going through LSTM one-by-one. Experimental results on two benchmarks demonstrate that the proposed method improves group activity recognition performance compared to the state-of-the-arts. Rui Yan 0010, Jinhui Tang 0001, Xiangbo Shu, Zechao Li, Qi Tian 0001 |
ACM Multimedia | 1 |
| 2018 | A Feature Selection Method for Projection Twin Support Vector Machine
Rui Yan 0010, Qiaolin Ye, Liyan Zhang 0001, Xiangbo Shu |
Neural Process. Lett. | 1 |