EDBT 2026 Demo / reviewers in the wild / expert
Yan Song 0005
dblp:09/1398-5
· DBLP profile ↗
20ranked-venue papers
2as first author
10since 2021 · last 2024
0000-0001-8431-7037ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Dilation-erosion for single-frame supervised temporal action localization
Yan Song 0005, Fanming Wang, Yang Zhao 0035, Xiangbo Shu, Yan Rui |
Multim. Tools Appl. | 2 |
| 2024 | Spatiotemporal Decouple-and-Squeeze Contrastive Learning for Semisupervised Skeleton-Based Action RecognitionabstractContrastive learning has been successfully leveraged to learn action representations for addressing the problem of semisupervised skeleton-based action recognition. However, most contrastive learning-based methods only contrast global features mixing spatiotemporal information, which confuses the spatial- and temporal-specific information reflecting different semantic at the frame level and joint level. Thus, we propose a novel spatiotemporal decouple-and-squeeze contrastive learning (SDS-CL) framework to comprehensively learn more abundant representations of skeleton-based actions by jointly contrasting spatial-squeezing features, temporal-squeezing features, and global features. In SDS-CL, we design a new spatiotemporal-decoupling intra-inter attention (SIIA) mechanism to obtain the spatiotemporal-decoupling attentive features for capturing spatiotemporal specific information by calculating spatial- and temporal-decoupling intra-attention maps among joint/motion features, as well as spatial- and temporal-decoupling inter-attention maps between joint and motion features. Moreover, we present a new spatial-squeezing temporal-contrasting loss (STL), a new temporal-squeezing spatial-contrasting loss (TSL), and the global-contrasting loss (GL) to contrast the spatial-squeezing joint and motion features at the frame level, temporal-squeezing joint and motion features at the joint level, as well as global joint and motion features at the skeleton level. Extensive experimental results on four public datasets show that the proposed SDS-CL achieves performance gains compared with other competitive methods. Binqian Xu, Xiangbo Shu, Jiachao Zhang, Guangzhao Dai, Yan Song 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Com-STAL: Compositional Spatio-Temporal Action LocalizationabstractSpatio-temporal action localization aims to locate the spatial and temporal positions of actors and classify their actions. However, prior research overlooks the fact that human actions often interact with novel objects in real-world scenarios, which neglects the various combinations of action-object, and considerably limits the generalization of the developed models. In this paper, we study the action-object combinations by researching multi-modal vision information of them. To this end, we propose a novel compositional spatio-temporal action localization (Com-STAL) task, which features non-overlapping action-object combinations in their training and test sets. Based on this, we construct a compositional action localization dataset (Com-AD). Beyond that, we propose a simple yet effective framework, Instance-Centric Interaction Network (ICIN), to reduce invalid induction biases within the visual modality and alleviate the combined distribution bias issue by leveraging additional modal information. The extensive experiment results on Com-AD demonstrate superior action localization performance of ICIN. Shaomeng Wang, Rui Yan 0010, Guangzhao Dai, Yan Song 0005, Xiangbo Shu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Turning to a Teacher for Timestamp Supervised Temporal Action SegmentationabstractTemporal action segmentation in videos has drawn much attention recently. Timestamp supervision is a cost-effective way for this task. To obtain more information to optimize the model, the existing method generated pseudo frame-wise labels iteratively based on the output of a segmentation model and the timestamp annotations. However, this practice may introduce noise and oscillation during the training, and lead to performance degeneration. To address this problem, we propose a new framework for timestamp supervised temporal action segmentation by introducing a teacher model parallel to the segmentation model to help stabilize the process of model optimization. The teacher model can be seen as an ensemble of the segmentation model, which helps to suppress the noise and to improve the stability of pseudo labels. We further introduce a segmentally smoothing loss, which is more focused and cohesive, to enforce the smooth transition of the predicted probabilities within action instances. The experiments on three datasets show that our method outperforms the state-of-the-art method and performs comparably against the fully-supervised methods at a much lower annotation cost. Yang Zhao 0035, Yan Song 0005 |
ICME | 2 |
| 2022 | Spatiotemporal Perturbation Based Dynamic Consistency for Semi-supervised Temporal Action Detection
Yan Song 0005, Rui Yan 0010, Xiangbo Shu |
MMM (1) | 2 |
| 2022 | Progressive enhancement network with pseudo labels for weakly supervised temporal action localization
Yan Song 0005, Xiangbo Shu |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Skip-attention encoder-decoder framework for human motion prediction
Xiangbo Shu, Rui Yan 0010, Jiachao Zhang, Yan Song 0005 |
Multim. Syst. | 5 |
| 2022 | Expansion-Squeeze-Excitation Fusion Network for Elderly Activity RecognitionabstractThis work focuses on the task of elderly activity recognition, which is a challenging task due to the existence of individual actions and human-object interactions in elderly activities. Thus, we attempt to effectively aggregate the discriminative information of actions and interactions from both RGB videos and skeleton sequences by attentively fusing multi-modal features. Recently, some nonlinear multi-modal fusion approaches are proposed by utilizing nonlinear attention mechanism that is extended from Squeeze-and-Excitation Networks (SENet). Inspired by this, we propose a novel Expansion-Squeeze-Excitation Fusion Network (ESE-FN) to effectively address the problem of elderly activity recognition, which learns modal and channel-wise Expansion-Squeeze-Excitation (ESE) attentions for attentively fusing the multi-modal features in the modal and channel-wise ways. Specifically, ESE-FN firstly implements the modal-wise fusion with the Modal-wise ESE Attention (M-ESEA) to aggregate discriminative information in modal-wise way, and then implements the channel-wise fusion with the Channel-wise ESE Attention (C-ESEA) to aggregate the multi-channel discriminative information in channel-wise way (referring toFigure 1). Furthermore, we design a new Multi-modal Loss (ML) to keep the consistency between the single-modal features and the fused multi-modal features by adding the penalty of difference between the minimum prediction losses on single modalities and the prediction loss on the fused modality. Finally, we conduct experiments on a largest-scale elderly activity dataset, i.e., ETRI-Activity3D (including 110,000+ videos, and 50+ categories), to demonstrate that the proposed ESE-FN achieves the best accuracy compared with the state-of-the-art methods. In addition, more extensive experimental results show that the proposed ESE-FN is also comparable to the other methods in terms of normal action recognition task. Xiangbo Shu, Jiawen Yang, Rui Yan 0010, Yan Song 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | X-Invariant Contrastive Augmentation and Representation Learning for Semi-Supervised Skeleton-Based Action RecognitionabstractSemi-supervised skeleton-based action recognition is a challenging problem due to insufficient labeled data. For addressing this problem, some representative methods leverage contrastive learning to obtain more features from the pre-augmented skeleton actions. Such methods usually adopt a two-stage way: first randomly augment samples, and then learn their representations via contrastive learning. Since skeleton samples have already been randomly augmented, the representation ability of the subsequent contrastive learning is limited due to the inconsistency between the augmentations and representations. Thus, we propose a novel X-invariant Contrastive Augmentation and Representation learning (X-CAR) framework to thoroughly obtain rotate-shear-scale (X for short) invariant features by learning augmentations and representations of skeleton sequences in a one-stage way. In X-CAR, a new Adaptive-combination Augmentation (AA) mechanism is designed to rotate, shear, and scale the skeletons by learnable controlling factors in an adaptive way rather than a random way. Here, such controlling factors are also learned in the whole contrastive learning process, which can facilitate the consistency between the learned augmentations and representations of skeleton sequences. In addition, we relax the pre-definition of positive and negative samples to avoid the confusing allocation of ambiguous samples, and present a new Pull-Push Contrastive Loss (PPCL) to pull the augmenting skeleton close to the original skeleton, while push far away from the other skeletons. Experimental results on both NTU RGB+D and North-Western UCLA datasets show that the proposed X-CAR achieves better accuracy compared with other competitive methods in the semi-supervised scenario. Binqian Xu, Xiangbo Shu, Yan Song 0005 |
IEEE Trans. Image Process. | 3 |
| 2021 | GRNet: Graph-based remodeling network for multi-view semi-supervised classification
Zhifan Zhu 0001, Yan Song 0005, Hai-juan Fu |
Pattern Recognit. Lett. | 3 |
| 2020 | Cross Fusion for Egocentric Interactive Action Recognition
Haiyu Jiang, Yan Song 0005, Xiangbo Shu |
MMM (1) | 2 |
| 2019 | Temporal Action Localization Based on Temporal Evolution Model and Multiple Instance Learning
Minglei Yang 0007, Yan Song 0005, Xiangbo Shu, Jinhui Tang 0001 |
MMM (2) | 2 |
| 2018 | Global and Local C3D Ensemble System for First Person Interactive Action Recognition
Lingling Fa, Yan Song 0005, Xiangbo Shu |
MMM (2) | 2 |
| 2018 | Effective Action Detection Using Temporal Context and Posterior Probability of Length
Yan Song 0005, Jinhui Tang 0001 |
MMM (2) | 2 |
| 2016 | A multi-phase sparse probability framework via entropy minimization for single sample face recognitionabstractIn this paper, we propose a robust probability based sparse method to solve single sample face recognition, which harvests the advantages of both local and global representation. Different from previous sparse representation methods that generate sparse coefficients by l1, we produce sparse class probability distribution by proposing a multi-phase sparse probability (MSP) framework. To create class probability distribution, we divide each face image into many local blocks and vote based on the classification results of all blocks. For classifying each block, we propose local similarity assumption that makes many conventional methods feasible to SSPP problem. Moreover, we also propose a heuristic multiphase class selection scheme to solve the entropy minimization problem, which finally provides a higher classification confidence from the global perspective. Experimental results on three popular databases show that our approach not only generalizes well to SSPP problem but also has strong robustness to expression, illumination, occlusion and time variation. Fan Liu 0003, Jinhui Tang 0001, Yan Song 0005, Qian Huang 0008, Feng Xu 0008 |
ICIP | 3 |
| 2016 | Local structure based multi-phase collaborative representation for face recognition with single sample per person
Fan Liu 0003, Jinhui Tang 0001, Yan Song 0005, Ye Bi, Sai Yang |
Inf. Sci. | 3 |
| 2015 | Describing Trajectory of Surface Patch for Human Action Recognition on RGB and Depth VideosabstractThis letter proposes a new feature describing the trajectories of surface patches (ToSP) on human bodies for action recognition by a novel scheme of utilizing RGB and depth videos. RGB data contains appearance information by which we track specific patches on body surfaces while depth data contains spatial information by which we describe surface patches. Specifically, we use spatial-temporal interest points as initial points to track in two directions. A ToSP is extracted by keeping the neighborhood in point cloud of each point on the trajectory. By using the temporal pyramid, ToSPs are matched on several levels based on the surface feature extracted from ToSP segments. The proposed feature captures both the shape and position variations of surface patches, thus it has the advantages of trajectories and local spatial-temporal features. The experiment results show that the proposed feature outperforms the existing trajectories based features and depth features. Yan Song 0005, Jinhui Tang 0001 |
IEEE Signal Process. Lett. | 1 |
| 2015 | Local Structure-Based Sparse Representation for Face RecognitionabstractThis article presents a simple yet effective face recognition method, called local structure-based sparse representation classification (LS_SRC). Motivated by the “divide-and-conquer” strategy, we first divide the face into local blocks and classify each local block, then integrate all the classification results to make the final decision. To classify each local block, we further divide each block into several overlapped local patches and assume that these local patches lie in a linear subspace. This subspace assumption reflects the local structure relationship of the overlapped patches, making sparse representation-based classification (SRC) feasible even when encountering the single-sample-per-person (SSPP) problem. To lighten the computing burden of LS_SRC, we further propose the local structure-based collaborative representation classification (LS_CRC). Moreover, the performance of LS_SRC and LS_CRC can be further improved by using the confusion matrix of the classifier. Experimental results on four public face databases show that our methods not only generalize well to SSPP problem but also have strong robustness to occlusion; little pose variation; and the variations of expression, illumination, and time. Fan Liu 0003, Jinhui Tang 0001, Yan Song 0005, Liyan Zhang 0001, Zhenmin Tang |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2014 | Local structure based sparse representation for face recognition with single sample per personabstractIn this paper, we propose local structure based sparse representation classification (LS SRC) to solve single sample per person (SSPP) problem. By adopting the “divide-conquer-aggregate” strategy, we successfully alleviate the dilemma of high data dimensionality and small samples, where we first divide the face into local blocks, and classify each local block, and then integrate all the classification results by voting. For each block, we further divide it into overlapped patches and assume that these patches lie in a linear subspace. This subspace assumption reflects local structure relationship of the overlapped patches and makes SRC feasible for SSPP problem. To lighten the computing burden, we further propose local structure based collaborative representation classification (LS CRC). Experimental results on three public face databases show that our methods not only generalize well to SSPP problem but also have strong robustness to expression, illumination, little pose variation, occlusion and time variation. Fan Liu 0003, Jinhui Tang 0001, Yan Song 0005, Xinguang Xiang, Zhenmin Tang |
ICIP | 3 |
| 2014 | Body Surface Context: A New Robust Feature for Action Recognition From Depth VideosabstractHuman action recognition in videos is useful for many applications. However, there still exist huge challenges in real applications due to the variations in the appearance, lighting condition and viewing angle, of the subjects. In this consideration, depth data have advantages over red, green, blue (RGB) data because of their spatial information about the distance between object and viewpoint. Unlike existing works, we utilize the 3-D point cloud, which contains points in the 3-D real-world coordinate system to represent the external surface of human body. Specifically, we propose a new robust feature, the body surface context (BSC), by describing the distribution of relative locations of the neighbors for a reference point in the point cloud in a compact and descriptive way. The BSC encodes the cylindrical angular of the difference vector based on the characteristics of human body, which increases the descriptiveness and discriminability of the feature. As the BSC is an approximate object-centered feature, it is robust to transformations including translations and rotations, which are very common in real applications. Furthermore, we propose three schemes to represent human actions based on the new feature, including the skeleton-based scheme, the random-reference-point scheme, and the spatial-temporal scheme. In addition, to evaluate the proposed feature, we construct a human action dataset by a depth camera. Experiments on three datasets demonstrate that the proposed feature outperforms RGB-based features and other existing depth-based features, which validates that the BSC feature is promising in the field of human action recognition. Yan Song 0005, Jinhui Tang 0001, Fan Liu 0003, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |