VLDB 2026 Research / reviewers in the wild / expert
Yang Chen 0039
dblp:48/4792-39
· DBLP profile ↗
10ranked-venue papers
7as first author
10since 2021 · last 2026
0009-0008-5752-1639ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STAR++: Region-Aware Conditional Semantics via Interpretable Side Information for Zero-Shot Skeleton Action RecognitionabstractZero-shot skeleton action recognition endeavors to classify novel action categories by transferring previously learned seen skeleton-semantic priors to unseen categories. However, current methods struggle to distinguish highly similar action categories, primarily due to the coarse-grained cross-modal alignment and non-discriminative representation space. To address these issues, we proposeSTAR++, a novel framework that aligns skeleton and semantics in a fine-grained and conditional manner. The key idea is to first establish region-level correspondences between body parts and semantic cues, and then utilize these local alignments to inform a global alignment process. This design is inspired by human visual cognition, which first attends to crucial local details before perceiving the broader scene. Concretely, we refine both skeleton and semantic representations with a dual-prompt attention mechanism driven by the structural decomposition of the human body and side information generated by a large language model (LLM). This encourages skeleton representations to be more compact within each class and semantic embeddings to be more separable across classes, which helps resolve ambiguity between highly similar actions and provides better interpretability of how unseen actions are perceived. Furthermore, we construct a region-aware holistic fusion module that aggregates these fine-grained features into a unified representation, yielding more discriminative holistic representations. Finally, the global alignment is conditioned on region-aware semantics feedback derived from fine-grained alignment, forming a conditional process that achieves more effective cross-modal alignment. Extensive experiments on four mainstream benchmarks demonstrate that our method achieves state-of-the-art performance in the zero-shot learning (ZSL) and generalized zero-shot learning (GZSL) settings. Yang Chen 0039, Jingcai Guo, Miaoge Li, Zhijie Rao, Song Guo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Neuron: Learning Context-Aware Evolving Representations for Zero-Shot Skeleton Action RecognitionabstractZero-shot skeleton action recognition is a non-trivial task that requires robust unseen generalization with prior knowledge from only seen classes and shared semantics. Existing methods typically build the skeleton-semantics interactions by uncontrollable mappings and conspicuous representations, thereby can hardly capture the intricate and fine-grained relationship for effective cross-modal transferability. To address these issues, we propose a novel dyNamically Evolving dUal skeleton-semantic syneRgistic framework with the guidance of cOntext-aware side informatioN (dubbed Neuron), to explore more fine-grained cross-modal correspondence from micro to macro perspectives at both spatial and temporal levels, respectively. Concretely, 1) we first construct the spatial-temporal evolving micro-prototypes and integrate dynamic context-aware side information to capture the intricate and synergistic skeleton-semantic correlations step-by-step, progressively refining cross-model alignment; and 2) we introduce the spatial compression and temporal memory mechanisms to guide the growth of spatial-temporal micro-prototypes, enabling them to absorb structure-related spatial representations and regularity-dependent temporal patterns. Notably, such processes are analogous to the learning and growth of neurons, equipping the framework with the capacity to generalize to novel unseen action categories. Extensive experiments on various benchmark datasets demonstrated the superiority of the proposed method1. Yang Chen 0039, Jingcai Guo, Song Guo 0001, Dacheng Tao |
CVPR | 1 |
| 2025 | Exploring Transferable Homogenous Groups for Compositional Zero-Shot LearningabstractConditional dependency present one of the trickiest problems in Compositional Zero-Shot Learning, leading to significant property variations of the same state (object) across different objects (states). To address this problem, existing approaches often adopt either all-to-one or one-to-one representation paradigms. However, these extremes create an imbalance in the seesaw between transferability and discriminability, favoring one at the expense of the other. Comparatively, humans are adept at analogizing and reasoning in a hierarchical clustering manner, intuitively grouping categories with similar properties to form cohesive concepts. Motivated by this, we propose Homogeneous Group Representation Learning (HGRL), a new perspective formulates state (object) representation learning as multiple homogeneous sub-group representation learning. HGRL seeks to achieve a balance between semantic transferability and discriminability by adaptively discovering and aggregating categories with shared properties, learning distributed group centers that retain group-specific discriminative features. Our method integrates three core components designed to simultaneously enhance both the visual and prompt representation capabilities of the model. Extensive experiments on three benchmark datasets validate the effectiveness of our method. Code is available at https://github.com/zjrao/HGRL. Zhijie Rao, Jingcai Guo, Miaoge Li, Yang Chen 0039, Mengzhu Wang |
IJCAI | 4 |
| 2025 | Multi-Level Skeleton Self-Supervised Learning: Enhancing 3D action representation learning with Large Multimodal Models
Yang Chen 0039, Ling Wang 0013, Rui Huang 0008, Hong Cheng 0002 |
Knowl. Based Syst. | 2 |
| 2025 | Enhancing Skeleton-Based Action Recognition With Language Descriptions From Pre-Trained Large Multimodal ModelsabstractSkeleton data has become popular in human action recognition because of its efficacy in capturing human motion patterns while mitigating the influence of environmental noise. However, overlooking critical action-related environmental descriptors presents challenges in distinguishing actions characterized by similar body movements. To address this limitation, we propose a novel framework that integrates skeleton data with language descriptions to easily capture essential environmental information for fine-grained action recognition while maintaining the robustness of skeleton-based methods. We first develop a Language Environment Description Generation (LEDG) module that utilizes the open-world understanding ability of Large Multimodal Models to generate instance-level action-related language environment descriptions without the need to train additional modules. Then, we introduce a Skeleton-supported Environment Feature Extraction (SEFE) module that leverages the temporal dependency inherent in skeleton data to extract key semantic environmental features. Additionally, we propose an Entropy-based Feature Fusion (EFF) module to dynamically amalgamate complementary features from both skeleton and language domains. Experimental results demonstrate the superiority of our framework, which can improve the accuracy of existing skeleton-based action recognition methods and achieve state-of-the-art performance on four well-established skeleton-based action recognition benchmarks. Yang Chen 0039, Ling Wang 0013, Hong Cheng 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Vision-Language Meets the Skeleton: Progressively Distillation With Cross-Modal Knowledge for 3D Action Representation LearningabstractSkeleton-based action representation learning aims to interpret and understand human behaviors by encoding the skeleton sequences, which can be categorized into two primary training paradigms: supervised learning and self-supervised learning. However, the former one-hot classification requires labor-intensive predefined action categories annotations, while the latter involves skeleton transformations (e.g., cropping) in the pretext tasks that may impair the skeleton structure. To address these challenges, we introduce a novel skeleton-based training framework (C$^{2}$VL) based onCross-modalContrastive learning that uses the progressive distillation to learn task-agnostic human skeleton action representation from theVision-Language knowledge prompts. Specifically, we establish the vision-language action concept space through vision-language knowledge prompts generated by pre-trained large multimodal models (LMMs), which enrich the fine-grained details that the skeleton action space lacks. Moreover, we propose the intra-modal self-similarity and inter-modal cross-consistency softened targets in the cross-modal representation learning process to progressively control and guide the degree of pulling vision-language knowledge prompts and corresponding skeletons closer. These soft instance discrimination and self-knowledge distillation strategies contribute to the learning of better skeleton-based action representations from the noisy skeleton-vision-language pairs. During the inference phase, our method requires only the skeleton data as the input for action recognition and no longer for vision-language prompts. Extensive experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD datasets demonstrate that our method outperforms the previous methods and achieves state-of-the-art results. Yang Chen 0039, Junfeng Fu, Ling Wang 0013, Jingcai Guo, Hong Cheng 0002 |
IEEE Trans. Multim. | 1 |
| 2024 | Fine-Grained Side Information Guided Dual-Prompts for Zero-Shot Skeleton Action RecognitionabstractSkeleton-based zero-shot action recognition aims to recognize unknown human actions based on the learned priors of the known skeleton-based actions and a semantic descriptor space shared by both known and unknown categories. However, previous works mainly focus on establishing the bridges between the known skeleton representation space and semantic descriptions space at the coarse-grained level for recognizing unknown action categories, ignoring the fine-grained alignment of these two spaces, resulting in suboptimal performance in distinguishing high-similarity action categories. To address these challenges, we propose a novel method via Side information and dual-prompTs learning for skeleton-based zero-shot Action Recognition (STAR) at the fine-grained level. Specifically, 1) we decompose the skeleton into several parts based on its topology structure and introduce the side information concerning multi-part descriptions of human body movements for alignment between the skeleton and the semantic space at the fine-grained level; 2) we design the visual-attribute and semantic-part prompts to improve the intra-class compactness within the skeleton space and inter-class separability within the semantic space, respectively, to distinguish the high-similarity actions. Extensive experiments show that our method achieves state-of-the-art performance in ZSL and GZSL settings on NTU RGB+D, NTU RGB+D 120, and PKU-MMD datasets. Yang Chen 0039, Jingcai Guo, Xiaocheng Lu, Ling Wang 0013 |
ACM Multimedia | 1 |
| 2024 | Spatio-temporal features for fast early warning of unplanned self-extubation in ICU
Yang Chen 0039, Ling Wang 0013, Guorong Wang, MingFang Xiang, Dekun Hu, Hong Cheng 0002 |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | PLFormer: Prompt Learning for Early Warning of Unplanned Extubation in ICUabstractPatients’ Unplanned Extubation (UEX) behaviors in ICU have adverse effects on their postoperative recovery. Therefore, it is necessary to design a early warning systems to detect UEX tendency. However, the fineness and rapidity of UEX behaviors, coupled with the complexity of the ICU environment, renders the utilization of RGB monitory videos for early warning extremely challenging. To address the aforementioned challenges, we propose a PLFormer to make early warning of UEX behaviors in ICU by using the prompt learning approach. Specifically, we provide click prompts to the Track Anything model (TAM) with the ability to segment and track patient regions, producing a mask sequence. Then we introduce the Prompt-Guided Adaptive Fusion (PAF) module, which utilizes mask prompts to guide the model’s attention towards the dynamically active region at both global and local levels. Subsequently, we incorporate the ST-Transformer to delve deeply into the spatial fine-grained representation and long-term temporal dependency properties of UEX behaviors. Experimental results demonstrate that our PLFormer achieves state-of-the-art performance on an ICU monitory dataset. Yang Chen 0039, Hong Cheng 0002, Ling Wang 0013 |
BIBM | 1 |
| 2023 | Multi-view graph convolution network for the recognition of human action with spatial and temporal occlusion problems
Yang Chen 0039, Ling Wang 0013, Dekun Hu, Hong Cheng 0002 |
J. Vis. Commun. Image Represent. | 1 |