Hongjun Li 0003

dblp:28/6464-3 · DBLP profile ↗
← Back
30ranked-venue papers
20as first author
28since 2021 · last 2026
0000-0001-7500-4979ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 15 first-author · 22 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Star-Shaped Multi-Person Interaction Graph Model for Group Skeleton-Based Action Recognition
abstract
Skeleton-based human action recognition has attracted increasing attention in recent years. However, most existing methods focus on single-person scenarios and struggle with complex behaviors in multi-person groups. In particular, they lack the capability to automatically identify and model core person. To address these challenges, this paper proposes a star-shaped group interaction model for skeleton-based action recognition. Firstly, the character importance scoring system analyzes both individual and group aspects: it evaluates each person's individual importance based on motion intensity and motion complexity, and assesses their significance within the group using centrality and interactivity. This process enables accurate identification of the core person in the video. Secondly, a core-star interaction graph is constructed with the core person as the center node and other individuals as peripheral nodes. The relationships among individuals are categorized into self-connections, centripetal connections, and centrifugal connections. For each type of connection, we design differentiated data augmentation strategies to fully exploit diverse action and interaction features. Finally, the structured skeleton data is fed into the star-shaped spatio-temporal graph convolutional network for efficient feature extraction and action classification. Experiments on several public benchmark datasets demonstrate that our method achieves state-of-the-art performance, achieving accuracies of 79.1%, 96.1%, and 93.1% on the NBA, Volleyball, and Volleyball-weak datasets, respectively.
Hongjun Li 0003, Yuehan Jiang, Liping Teng
IEEE Trans. Image Process.1
2025 Enhancing Text-to-Image Diffusion with Intent-Aware Semantic Alignment from Simple Prompts
Yangyu Liu, Mingzi Chen, Hongjun Li 0003
IEEE Big Data3
2025 MPAF-SECL: Multi-path Adaptive Fusion with Semantic-Enhanced Cross-Modal Learning for Skeleton-Based Action Recognition
Tian Bai 0013, Zhengjie Ni, Hongjun Li 0003
PRCV (7)3
2025 Fusing Rigid Skeletal Nodes to Graph Convolutional Networks for Fuzzy Action Recognition
Yuehan Jiang, Hongjun Li 0003
PRCV (7)2
2025 A novel spatio-temporal memory network for video anomaly detection
Hongjun Li 0003
Multim. Tools Appl.1
2025 Adaptive loitering anomaly detection based on motion states
Hongjun Li 0003, Xiezhou Huang
Multim. Tools Appl.1
2025 SACM: spatial attributes of complex movements in multi-object tracking
Hongjun Li 0003, Jiaxin Li 0003, Xiaohu Sun
Multim. Tools Appl.1
2025 A multi-memory-augmented network with a curvy metric method for video anomaly detection
Hongjun Li 0003
Neural Networks1
2025 Action-Responsive Contrastive Network for Fine-Grained Skeleton-Based Action Recognition
abstract
Currently, fine-grained skeleton action recognition based on graph convolutional networks (GCNs) has become an important research focus. Fine-grained action recognition refers to the accurate recognition of subtle, complex or detailed actions. This task is particularly challenging due to the limited appearance information in skeleton data and the limitations of predefined single-topology skeleton structures. To address these challenges, we propose an action-responsive contrastive network (ARCN). The network consists of two main components: an action-responsive graph convolutional network (ARGCN) with enhanced skeleton topology and a fine-grained action comparator (FAC) that uses feature contrastive learning to explore the latent space of motion features. The ARGCN contains two specialized modules: the action-responsive topology (ART) module, which captures important motion features through the learned action-specific topology structure matrix and multiscale temporal features; and the action-responsive attention (ARA) module, which learns complex spatiotemporal skeleton attention information. These modules jointly generate a multichannel cross-temporal dynamic skeleton joint attention topology map tailored for the specific action being analysed. To further clarify the fine-grained action feature differences, the FAC is integrated in some stages of the ARGCN. The FAC performs spatiotemporal decoupling of feature maps, classifies and contrasts similar and different fine-grained motion features, and builds a learnable latent space for fine-grained motion, thereby improving classification performance. Our model is evaluated on six public datasets: NTU RGB+D, NTU RGB+D 120, NW-UCLA, UAV-Human, Finegym, and Diving48. It achieves 91.2% accuracy on the NTU RGB+D 120 dataset X-Set, 97.2% accuracy on the NW-UCLA dataset, 44.6% accuracy on the UAV-Human dataset CSv1, 72.0% accuracy on the UAV-Human dataset CSv2, 95.3% accuracy on the Finegym dataset, and 54.3% accuracy on the Diving48 dataset, which are competitive results compared with the state-of-the-art methods.
Hongjun Li 0003, Tian Bai 0013
IEEE Trans. Multim.1
2025 Detecting Adversarial Attacks Based on Tracking Differences in Frequency Bands
abstract
Like deep neural network (DNN)-based classifiers, DNN-based trackers are also vulnerable to adversarial attacks that degrade the tracking performance by adding adversarial perturbations to the input videos. This paper proposes a detection method for the first time to assist the tracker in detecting adversarial attacks. The adversarial perturbations in the visual object tracking task are invisible but are effective at attacking trackers. This naturally creates challenges in detecting attacks in the spatial pixel domain. To this end, we innovatively transfer the detection of adversarial attacks from the spatial domain to the frequency domain. Specifically, we first theoretically prove that the perturbations are added mainly to the high-frequency band of the video. Then, from the empirical studies, we conclude that the low-frequency band contributes most to the tracking performance and is most robust against adversarial attacks. According to the theoretical proof and empirical conclusion, we finally design an unsupervised adversarial detection framework, which mainly contains a frequency decomposition module (FDM), a target tracker (TT) with its mirror tracker (MT), and a discriminant module (DM). For an input video, the TT is fed the full-frequency video, whereas the MT takes as input the low-frequency video that is decomposed by the FDM. The DM discriminates the input video as adversarial or natural by comparing the racking performance differences between the two trackers. The whole detection process is performed along with the tracking phase, and all the modules in the framework require no training on adversarial examples. Extensive experiments demonstrate that our adversarial detection framework can effectively detect mainstream adversarial attacks in the tracking field. It can also be flexibly integrated with many trackers, including anchor-based and anchor-free trackers. More importantly, the trackers integrated with the detection framework can still maintain near-original tracking performance.
Hongjun Li 0003, Guoan Zhang
IEEE Trans. Multim.2
2024 FOAD: a novel video anomaly detection focusing on objects
Hongjun Li 0003, Xiezhou Huang, Yunlong Du
Multim. Tools Appl.1
2024 MTM-net: a multidimensional two-stage memory-guided network for vedio abnormal detection
Hongjun Li 0003, Xiaohu Sun
Multim. Tools Appl.1
2024 TMTB: Transformer based multi-task branching multi-object tracking algorithm for wide-view scenes
Hongjun Li 0003, Jiaxin Li 0003
Multim. Tools Appl.1
2024 Grey-adversary perceptual network for anomaly detection
Hongjun Li 0003, Guoan Zhang
Multim. Tools Appl.2
2024 Appearance-motion heterogeneous networks for video anomaly detection
Hongjun Li 0003, Xiaohu Sun
Multim. Tools Appl.1
2024 Channel based approach via faster dual prediction network for video anomaly detection
Hongjun Li 0003, Xulin Shen, Xiaohu Sun
Multim. Tools Appl.1
2024 Cross-modality integration framework with prediction, perception and discrimination for video anomaly detection
Hongjun Li 0003, Guoan Zhang
Neural Networks2
2024 An informative dual ForkNet for video anomaly detection
Hongjun Li 0003
Neural Networks1
2023 Transformer-Based Multi-object Tracking in Unmanned Aerial Vehicles
Jiaxin Li 0003, Hongjun Li 0003
PRCV (6)2
2023 Future frame prediction based on generative assistant discriminative network for anomaly detection
Hongjun Li 0003, Guoan Zhang
Appl. Intell.2
2023 Multi-memory video anomaly detection based on scene object distribution
Hongjun Li 0003, Xiaohu Sun
Multim. Tools Appl.1
2023 MPAT: multi-path attention temporal method for video anomaly detection
Hongjun Li 0003, Xiaohu Sun, Xulin Shen, Zhengguang Xie
Multim. Tools Appl.1
2023 Video anomaly detection based on scene classification
Hongjun Li 0003, Xulin Shen, Xiaohu Sun
Multim. Tools Appl.1
2023 HN-MUM: heterogeneous video anomaly detection network with multi-united-memory module
Hongjun Li 0003, Jiaxin Li 0003
Multim. Tools Appl.1
2023 Regression-Selective Feature-Adaptive Tracker for Visual Object Tracking
abstract
As a challenging visual task, visual object tracking has recently been composed of the classification and regression subtasks. The anchor-free regression network gets rid of the dependence on the anchors, but the redundant range makes it usually regress some samples involving non-target information. Evenly dividing a target by the regular receptive field often causes ambiguous target localization. To address these issues, we propose a regression-selective feature-adaptive tracker (RSFA), where the regression-selective subnetwork can not only free the regression task from anchors, but can also select more effective regression samples using the refined criterion. The proposed feature-adaptive strategy concentrates the classification subnetwork on target feature extraction via adaptively modifying the receptive field, and the attached centrality branch offers a correction for target localization by exploiting the spatial information. Additionally, the designed online update mechanism realizes the tracker's online optimization, improving robustness against target deformation. Extensive experiments are conducted on challenging benchmarks, including GOT10 K, OTB2015, UAV123, NFS, VOT2018, VOT2019 and VOT2020-ST. Our tracker achieves satisfactory tracking results, and the evaluations of its tracking performance rank first or second in comparison with the state-of-the-art tracking algorithms.
Ze Zhou 0002, Quan-Sen Sun, Hongjun Li 0003, Zhenwen Ren
IEEE Trans. Multim.3
2022 A novel vertical-cross-horizontal network
Ze Zhou 0002, Hongjun Li 0003, Zhengguang Xie, Guoan Zhang
Multim. Tools Appl.3
2022 A near effective and efficient model in recognition
Hongjun Li 0003, Ze Zhou 0002, Ching Y. Suen
Pattern Recognit.1
2021 Fall detection based on fused saliency maps
Hongjun Li 0003, Yupeng Ding
Multim. Tools Appl.1
2016 A novel Non-local means image denoising method based on grey theory
Hongjun Li 0003, Ching Y. Suen
Pattern Recognit.1
2016 Robust face recognition based on dynamic rank representation
Hongjun Li 0003, Ching Y. Suen
Pattern Recognit.1