Yan Gao 0025

dblp:46/3479-25 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
8since 2021 · last 2025
0009-0006-6216-3165ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 WM-SORT: Modeling multi-object tracking motion patterns in real world
Yizhuo Jiang, Weiyu Zhao, Yan Gao 0025, Xinhang Niu, Jie Li 0001, Xinbo Gao 0001
Neurocomputing3
2025 DETrack: Depth information is predictable for tracking
Weiyu Zhao, Yizhuo Jiang, Yan Gao 0025, Jie Li 0001, Xinbo Gao 0001
Neurocomputing3
2025 Language Knowledge-Assisted Representation Learning for Skeleton-Based Action Recognition
abstract
How humans understand and recognize the actions of others is a complex neuroscientific problem that involves a combination of cognitive mechanisms and neural networks. Research has shown that humans have brain areas that recognize actions that process top-down attentional information, such as the temporoparietal association area. Also, humans have brain regions dedicated to understanding the minds of others and analyzing their intentions, such as the medial prefrontal cortex of the temporal lobe. Skeleton-based action recognition creates mappings for the complex connections between the human skeleton movement patterns and behaviors. Although existing studies encoded meaningful node relationships and synthesized action representations for classification with good results, few of them considered incorporating a priori knowledge to aid potential representation learning for better performance. LA-GCN proposes a graph convolution network using large-scale language models (LLM) knowledge assistance. First, the LLM knowledge is mapped into a priori global relationship (GPR) topology and a priori category relationship (CPR) topology between nodes. The GPR guides the generation of new “bone” representations, aiming to emphasize essential node information from the data level. The CPR mapping simulates category prior knowledge in human brain regions, encoded by the PC-AC module and used to add additional supervision—forcing the model to learn class-distinguishable features. In addition, to improve information transfer efficiency in topology modeling, we propose multi-hop attention graph convolution. It aggregates each node's k-order neighbor simultaneously to speed up model convergence. LA-GCN reaches state-of-the-art on NTU RGB+D, NTU RGB+D 120, and NW-UCLA datasets.
Yan Gao 0025, Zheng Hui, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Multim.2
2025 An Information Compensation Framework for Zero-Shot Skeleton-Based Action Recognition
abstract
Zero-shot human skeleton-based action recognition aims to construct a model that can recognize actions outside the categories seen during training. Previous research has focused on aligning sequences' visual and semantic spatial distributions. However, these methods extract semantic features simply. They ignore that proper prompt design for rich and fine-grained action cues can provide robust representation space clustering. In order to alleviate the problem of insufficient information available for skeleton sequences, we design an information compensation learning framework from an information-theoretic perspective to improve zero-shot action recognition accuracy with a multi-granularity semantic interaction mechanism. Inspired by ensemble learning, we propose a multi-level alignment (MLA) approach to compensate information for action classes. MLA aligns multi-granularity embeddings with visual embedding through a multi-head scoring mechanism to distinguish semantically similar action names and visually similar actions. Furthermore, we introduce a new loss function sampling method to obtain a tight and robust representation. Finally, these multi-granularity semantic embeddings are synthesized to form a proper decision surface for classification. Significant action recognition performance is achieved when evaluated on the challenging NTU RGB+D, NTU RGB+D 120, and PKU-MMD benchmarks and validate that multi-granularity semantic features facilitate the differentiation of action clusters with similar visual features.
Yan Gao 0025, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Multim.2
2024 Multi-Scene Generalized Trajectory Global Graph Solver with Composite Nodes for Multiple Object Tracking
abstract
The global multi-object tracking (MOT) system can consider interaction, occlusion, and other ``visual blur'' scenarios to ensure effective object tracking in long videos. Among them, graph-based tracking-by-detection paradigms achieve surprising performance. However, their fully-connected nature poses storage space requirements that challenge algorithm handling long videos. Currently, commonly used methods are still generated trajectories by building one-forward associations across frames. Such matches produced under the guidance of first-order similarity information may not be optimal from a longer-time perspective. Moreover, they often lack an end-to-end scheme for correcting mismatches. This paper proposes the Composite Node Message Passing Network (CoNo-Link), a multi-scene generalized framework for modeling ultra-long frames information for association. CoNo-Link's solution is a low-storage overhead method for building constrained connected graphs. In addition to the previous method of treating objects as nodes, the network innovatively treats object trajectories as nodes for information interaction, improving the graph neural network's feature representation capability. Specifically, we formulate the graph-building problem as a top-k selection task for some reliable objects or trajectories. Our model can learn better predictions on longer-time scales by adding composite nodes. As a result, our method outperforms the state-of-the-art in several commonly used datasets.
Yan Gao 0025, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
AAAI1
2024 Trades++: Enhancing Multi-Object Tracking of Real Low Confidence Targets Using a Pyramid-Like Self-Attention Model
abstract
In reality, multi-object tracking (MOT) is used in a wide range of scenarios. Maintaining the motion trajectory of the target, especially in high-density pedestrian scenarios, is often difficult. The tracking quality of most multi-object trackers correlates strongly with the detector quality and they often ignore the low-scoring detection boxes obtained by the detector. In this paper, we propose a TraDeS-based method called TraDeS++ that enhances the detection features using a pyramid-like self-attention model, significantly reducing the model training time and achieving a reduction of half the training epochs. The second motivation is to focus on the association method. We use a two-stage matching strategy with GIoU constraints, effectively improving HOTA. Experimental results show that our component effectively improves the metrics of MOT, especially MOTA, HOTA, and IDF1. Competitive results are achieved on the popular MOT16 and MOT17 datasets.
Chenxin Wen, Yan Gao 0025, Jie Li 0001
ICASSP2
2024 BPMTrack: Multi-Object Tracking With Detection Box Application Pattern Mining
abstract
The key to multi-object tracking is its stability and the retention of identity information. A common problem with most detection-based approaches is trusting and using all the detector outputs for the association. However, some settings of detectors can affect stable long-range tracking. Based on the principle of reducing the association noise in the detection processing step, we propose a new framework, the Box application Pattern Mining Tracker (BPMTrack), to address this issue. Specifically, we worked on three main aspects: output threshold, association strategy, and motion model. Due to the problem of inconsistency between classification scores and localization accuracy, we propose the Box Quality Estimation Network (BQENet) to predict the localization quality scores of all detections in the current frame, reserving high-quality boxes for the tracker. In addition, based on observations of intensive scenarios, we propose a simple and effective data association method, the Non-Maximum Suppression Integration (NMSI) matching strategy. It recovers the Non-Maximum Suppression (NMS) detection, inputs them into BQENet, and then performs hierarchical matching with reasonable control of box priority to alleviate the problem of absent objects caused by occlusion. Finally, we propose an improved Measurement Correct and Noise Scale (MCNS) Kalman algorithm to improve the prediction accuracy of object positions and, thus, the association quality. We performed an extensive ablation evaluation of the proposed framework to prove its effectiveness. Moreover, the three tracking benchmarks show our method's accuracy and long-distance performance.
Yan Gao 0025, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2022 An Object Point Set Inductive Tracker for Multi-Object Tracking and Segmentation
abstract
Multi-object tracking and segmentation (MOTS) is a derivative task of multi-object tracking (MOT). The new setting encourages the learning of more discriminative high-quality embeddings. In this paper, we focus on the problem of exploring the relationship between the segmenter and the tracker, and propose an efficient Object Point set Inductive Tracker (OPITrack) based on it. First, we discover that after a single attention layer, the high-dimensional, key point embedding will show feature averaging. To alleviate this phenomenon, we propose an embedding generalization training strategy for sparse training and dense testing. This strategy allows the network to increase randomness in training and encourages the tracker to learn more discriminative features. In addition, to learn the desired embedding space, we propose a general Trip-hard sample augmentation loss. The loss uses patches that are not distinguishable by the segmenter to join the feature learning and force the embedding network to learn the difference between false positives and true positives. Our method was validated on two MOTS benchmark datasets and achieved promising results. In addition, our OPITrack can achieve better performance for the raw model while costing less video memory (VRAM) at training time.
Yan Gao 0025, Yu Zheng 0006, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2020 CBFNet: Constraint balance factor for semantic segmentation
Yan Gao 0025, Jie Li 0001, Xinbo Gao 0001
Neurocomputing2