VLDB 2026 Research / reviewers in the wild / expert
Ling Wang 0013
dblp:45/6607-13
· DBLP profile ↗
11ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-1235-1241ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Unified Spatiotemporal Frequency Graph Neural Network for fMRI-based Brain Functional Connectivity AnalysisabstractAnalyzing functional connectivity patterns from resting-state functional magnetic resonance imaging (fMRI) requires unraveling its interrelations across spatial, temporal, and frequency domains. To comprehensively analyze four-dimensional (4D) fMRI data, we propose the Spatiotemporal Frequency Graph Neural Network (STFreqGNN), which processes dynamic heterogeneous graphs across spatial, temporal, and frequency domains using a transformer-style architecture. To reduce the complexity of multi-domain analysis with small sample sizes for fMRI datasets and ensure domainspecific interpretability, we introduce two structure-informed modules in the spatial and temporal domains to improve knowledge aggregation within each domain. Specifically, the plugin GNNs transmit information within the static homogeneous brain region graphs, and recurrent blocks aggregate features from heterogeneous nodes defined across different temporal windows. Additionally, we design cross-domain masked self-attention blocks to prevent attention captured by irrelevant or redundant token pairs, finally enabling efficient disease-specific feature learning. Experimental results on both public and in-house datasets suggest that the proposed method is not only superior to several state-of-the-art methods on fMRI-based classification but also preserves interpretation ability in all these domains. Yulang Huang, Zhiyuan Ding, Guokai Duan, Yan Liu 0054, Xiangzhu Zeng, Ling Wang 0013 |
ICASSP | 8 |
| 2025 | Multi-Level Skeleton Self-Supervised Learning: Enhancing 3D action representation learning with Large Multimodal Models
Yang Chen 0039, Ling Wang 0013, Rui Huang 0008, Hong Cheng 0002 |
Knowl. Based Syst. | 4 |
| 2025 | Enhancing Skeleton-Based Action Recognition With Language Descriptions From Pre-Trained Large Multimodal ModelsabstractSkeleton data has become popular in human action recognition because of its efficacy in capturing human motion patterns while mitigating the influence of environmental noise. However, overlooking critical action-related environmental descriptors presents challenges in distinguishing actions characterized by similar body movements. To address this limitation, we propose a novel framework that integrates skeleton data with language descriptions to easily capture essential environmental information for fine-grained action recognition while maintaining the robustness of skeleton-based methods. We first develop a Language Environment Description Generation (LEDG) module that utilizes the open-world understanding ability of Large Multimodal Models to generate instance-level action-related language environment descriptions without the need to train additional modules. Then, we introduce a Skeleton-supported Environment Feature Extraction (SEFE) module that leverages the temporal dependency inherent in skeleton data to extract key semantic environmental features. Additionally, we propose an Entropy-based Feature Fusion (EFF) module to dynamically amalgamate complementary features from both skeleton and language domains. Experimental results demonstrate the superiority of our framework, which can improve the accuracy of existing skeleton-based action recognition methods and achieve state-of-the-art performance on four well-established skeleton-based action recognition benchmarks. Yang Chen 0039, Ling Wang 0013, Hong Cheng 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Vision-Language Meets the Skeleton: Progressively Distillation With Cross-Modal Knowledge for 3D Action Representation LearningabstractSkeleton-based action representation learning aims to interpret and understand human behaviors by encoding the skeleton sequences, which can be categorized into two primary training paradigms: supervised learning and self-supervised learning. However, the former one-hot classification requires labor-intensive predefined action categories annotations, while the latter involves skeleton transformations (e.g., cropping) in the pretext tasks that may impair the skeleton structure. To address these challenges, we introduce a novel skeleton-based training framework (C$^{2}$VL) based onCross-modalContrastive learning that uses the progressive distillation to learn task-agnostic human skeleton action representation from theVision-Language knowledge prompts. Specifically, we establish the vision-language action concept space through vision-language knowledge prompts generated by pre-trained large multimodal models (LMMs), which enrich the fine-grained details that the skeleton action space lacks. Moreover, we propose the intra-modal self-similarity and inter-modal cross-consistency softened targets in the cross-modal representation learning process to progressively control and guide the degree of pulling vision-language knowledge prompts and corresponding skeletons closer. These soft instance discrimination and self-knowledge distillation strategies contribute to the learning of better skeleton-based action representations from the noisy skeleton-vision-language pairs. During the inference phase, our method requires only the skeleton data as the input for action recognition and no longer for vision-language prompts. Extensive experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD datasets demonstrate that our method outperforms the previous methods and achieves state-of-the-art results. Yang Chen 0039, Junfeng Fu, Ling Wang 0013, Jingcai Guo, Hong Cheng 0002 |
IEEE Trans. Multim. | 4 |
| 2024 | Fine-Grained Side Information Guided Dual-Prompts for Zero-Shot Skeleton Action RecognitionabstractSkeleton-based zero-shot action recognition aims to recognize unknown human actions based on the learned priors of the known skeleton-based actions and a semantic descriptor space shared by both known and unknown categories. However, previous works mainly focus on establishing the bridges between the known skeleton representation space and semantic descriptions space at the coarse-grained level for recognizing unknown action categories, ignoring the fine-grained alignment of these two spaces, resulting in suboptimal performance in distinguishing high-similarity action categories. To address these challenges, we propose a novel method via Side information and dual-prompTs learning for skeleton-based zero-shot Action Recognition (STAR) at the fine-grained level. Specifically, 1) we decompose the skeleton into several parts based on its topology structure and introduce the side information concerning multi-part descriptions of human body movements for alignment between the skeleton and the semantic space at the fine-grained level; 2) we design the visual-attribute and semantic-part prompts to improve the intra-class compactness within the skeleton space and inter-class separability within the semantic space, respectively, to distinguish the high-similarity actions. Extensive experiments show that our method achieves state-of-the-art performance in ZSL and GZSL settings on NTU RGB+D, NTU RGB+D 120, and PKU-MMD datasets. Yang Chen 0039, Jingcai Guo, Xiaocheng Lu, Ling Wang 0013 |
ACM Multimedia | 5 |
| 2024 | Spatio-temporal features for fast early warning of unplanned self-extubation in ICU
Yang Chen 0039, Ling Wang 0013, Guorong Wang, MingFang Xiang, Dekun Hu, Hong Cheng 0002 |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | PLFormer: Prompt Learning for Early Warning of Unplanned Extubation in ICUabstractPatients’ Unplanned Extubation (UEX) behaviors in ICU have adverse effects on their postoperative recovery. Therefore, it is necessary to design a early warning systems to detect UEX tendency. However, the fineness and rapidity of UEX behaviors, coupled with the complexity of the ICU environment, renders the utilization of RGB monitory videos for early warning extremely challenging. To address the aforementioned challenges, we propose a PLFormer to make early warning of UEX behaviors in ICU by using the prompt learning approach. Specifically, we provide click prompts to the Track Anything model (TAM) with the ability to segment and track patient regions, producing a mask sequence. Then we introduce the Prompt-Guided Adaptive Fusion (PAF) module, which utilizes mask prompts to guide the model’s attention towards the dynamically active region at both global and local levels. Subsequently, we incorporate the ST-Transformer to delve deeply into the spatial fine-grained representation and long-term temporal dependency properties of UEX behaviors. Experimental results demonstrate that our PLFormer achieves state-of-the-art performance on an ICU monitory dataset. Yang Chen 0039, Hong Cheng 0002, Ling Wang 0013 |
BIBM | 5 |
| 2023 | Multi-view graph convolution network for the recognition of human action with spatial and temporal occlusion problems
Yang Chen 0039, Ling Wang 0013, Dekun Hu, Hong Cheng 0002 |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Analysis on Teeth Occlusion Distribution Based on Segmentation and Registration AlgorithmabstractOcclusal contact status of teeth is a key indicator for orthodontic and periodontal disease treatment. Digital analysis includes tooth position recognition and occlusal contact distribution estimation. In this study, we propose a cascade two-stage point-wise network named Teeth Segmentation Network (TSegNet) based on self-attention mechanism to address teeth segmentation task. And a template-based registration method is proposed to analyze the status of teeth occlusion. In TSegNet, spatial and channel attention are used to improve the performance feature extraction. Template-registration-based occlusal distribution analysis method reduced the labeled number of training samples. To the best of our knowledge, it is the first study on occlusal contact analyzing by using computer-aided-diagnosis technique. Experiment results illustrate the effectiveness and robustness of our proposed method. Zihan Cao, Xinwu Sun, Gangyuan Chen, Yan Liu 0054, Xinggang Liu, Dongxiang Zheng, Ling Wang 0013 |
BIBM | 8 |
| 2018 | A set-to-set nearest neighbor approach for robust and efficient face recognition with image sets
Ling Wang 0013, Hong Cheng 0002, Zicheng Liu 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | A robust elastic net approach for feature learning
Ling Wang 0013, Hong Cheng 0002, Zicheng Liu 0001, Ce Zhu |
J. Vis. Commun. Image Represent. | 1 |