EDBT 2026 Demo / reviewers in the wild / expert
Jianyu Yang 0002
dblp:23/6903-2
· DBLP profile ↗
29ranked-venue papers
8as first author
14since 2021 · last 2024
0000-0002-0208-221XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A temporal densely connected recurrent network for event-based human pose estimation
Zhanpeng Shao, Wuzhen Wang, Jianyu Yang 0002, Youfu Li 0001 |
Pattern Recognit. | 5 |
| 2024 | You Will Never Walk Alone: One-Shot 3D Action Recognition With Point Cloud SequenceabstractIn this work, we pay the first effort to address one-shot 3D action recognition in point cloud sequence, without skeleton information. The main contribution lies in two folders. First, a novel one-shot classification approach that considers the feature distribution of 3D action is proposed. We find that, for different 3D actions their dimensional-wise feature distributions are generally in Gaussian form and similar action categories hold approximate feature distributions. Accordingly, K-nearest base classes’ mean value and covariance matrix information help to form one-shot novel class’s pseudo feature distribution. To alleviate the potential ambiguous problem within nearest neighbor search, we divide the base classes into subsets via C-means clustering to facilitate the similarity measure to novel class. Meanwhile, the feature distribution of base class’s whole set and subsets will be jointly considered for generating novel class’s pseudo feature distribution. Multi-dimensional Gaussian sampling is conducted on the acquired pseudo feature distribution for feature-level data augmentation, to make one-shot novel class “never walk alone” for leveraging classifier training. Secondly to better characterize fine-grained 3D action, a temporal attention method is proposed, via introducing vision Transformer (ViT) to capture action’s discriminative short-term motion pattern with densely sampled short-term 3DV (3D dynamic voxel) features along temporal dimension. Experiments on NTU RGB+D 120 and 60 verify superiority of our approach. It outperforms state-of-the-art skeleton-based methods by 13.9% at most. The source code is available athttps://github.com/Tong-XY/YNWA. Xingyu Tong 0002, Yang Xiao 0007, Jianyu Yang 0002, Zhiguo Cao 0001, Joey Tianyi Zhou, Junsong Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Beyond Pattern Variance: Unsupervised 3-D Action Representation Learning With Point Cloud SequenceabstractThis work pays the first research effort to address unsupervised 3-D action representation learning with point cloud sequence, which is different from existing unsupervised methods that rely on 3-D skeleton information. Our proposition is built on the state-of-the-art 3-D action descriptor 3-D dynamic voxel (3DV) with contrastive learning (CL). The 3DV can compress the point cloud sequence into a compact point cloud of 3-D motion information. Spatiotemporal data augmentations are conducted on it to drive CL. However, we find that existing CL methods (e.g., SimCLR or MoCo v2) often suffer from high pattern variance toward the augmented 3DV samples from the same action instance, that is, the augmented 3DV samples are still of high feature complementarity after CL, while the complementary discriminative clues within them have not been well exploited yet. To address this, a feature augmentation adapted CL (FACL) approach is proposed, which facilitates 3-D action representation via concerning the features from all augmented 3DV samples jointly, in spirit of feature augmentation. FACL runs in a global-local way: one branch learns global feature that involves the discriminative clues from the raw and augmented 3DV samples, and the other focuses on enhancing the discriminative power of local feature learned from each augmented 3DV sample. The global and local features are fused to characterize 3-D action jointly via concatenation. To fit FACL, a series of spatiotemporal data augmentation approaches is also studied on 3DV. Wide-range experiments verify the superiority of our unsupervised learning method for 3-D action feature learning. It outperforms the state-of-the-art skeleton-based counterparts by 6.4% and 3.6% with the cross-setup and cross-subject test settings on NTU RGB+D 120, respectively. The source code is available at https://github.com/tangent-T/FACL. Yang Xiao 0007, Yancheng Wang 0002, Jianyu Yang 0002, Zhiguo Cao 0001, Joey Tianyi Zhou, Junsong Yuan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Learning dynamic relationship between joints for 3D hand pose estimation from single depth map
Huiqin Xing, Jianyu Yang 0002, Yang Xiao 0007 |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | Learning full context feature for human motion prediction
Huiqin Xing, Yicong Zhou, Jianyu Yang 0002, Yang Xiao 0007 |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | Corner Detection via Scale-Space Behavior-Guided Trajectory TracingabstractExistingcurvature scale-space(CSS) methods detect corners by tracing the CSS trajectories from a determined high scale toward the lowest one. For those images with sophisticated details, such approach could often yield unsatisfactory corner detection results; i.e., miss-detectedtruecorners (false negatives) androundcorners (false positives). In this letter, these two fundamental problems are investigated. To tackle them, a novel CSS-based corner detector is proposed by incorporating our mathematically derived scale-space properties of the planar curves and corner points into the developed trajectory tracing algorithm, called thescale-space behavior-guided trajectory tracing(SBTT). In view of lacking a benchmark dataset with the ground truth, another contribution from our work is on the establishment of an augmented test image dataset, containing 147 test images with manually-labelled ground truth and their augmented images up to 62,328 images in total. Based on the ground truth, four commonly-used metrics are exploited to conduct corner detection performance evaluation. The obtained simulation results show that our proposed corner detector yields the highest F-score, when compared with that of nine state-of-the-art methods. Baojiang Zhong, Jianyu Yang 0002, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 3 |
| 2022 | MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in VideoabstractRecent transformer-based solutions have been introduced to estimate 3D human pose from 2D keypoint sequence by considering body joints among all frames globally to learn spatio-temporal correlation. We observe that the motions of different joints differ significantly. However, the previous methods cannot efficiently model the solid inter-frame correspondence of each joint, leading to insufficient learning of spatial-temporal correlation. We propose MixSTE (Mixed Spatio-Temporal Encoder), which has a temporal transformer block to separately model the temporal motion of each joint and a spatial transformer block to learn inter-joint spatial correlation. These two blocks are utilized alternately to obtain better spatio-temporal feature encoding. In addition, the network output is extended from the central frame to entire frames of the input video, thereby improving the coherence between the input and output sequences. Extensive experiments are conducted on three benchmarks (i.e. Human3.6M, MPI-INF-3DHP, and HumanEva). The results show that our model outperforms the state-of-the-art approach by 10.9% P-MPJPE and 7.6% MPJPE. The code is available at https://github.com/JinluZhang1126/MixSTE. Jinlu Zhang 0001, Zhigang Tu 0001, Jianyu Yang 0002, Yujin Chen, Junsong Yuan 0001 |
CVPR | 3 |
| 2022 | FastHand: Fast monocular hand pose estimation on embedded systems
Shan An, Xiajie Zhang, Haogang Zhu, Jianyu Yang 0002, Konstantinos A. Tsintotas |
J. Syst. Archit. | 5 |
| 2022 | Multi-stream feature refinement network for human object interaction detection
Zhanpeng Shao, Zhongyan Hu, Jianyu Yang 0002, Youfu Li 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2022 | Multibranch Adversarial Regression for Domain Adaptative Hand Pose EstimationabstractAlthough hand pose estimation has achieved a great success in recent years, there are still challenges with RGB-based estimation tasks, the most significant of which is the absence of labeled training data. At present, the synthetic dataset has plenty of images with accurate annotation, but the difference from real-world datasets affects generalization. Therefore, a transfer learning strategy, which tries to transfer knowledge from a labeled source domain to an unlabeled target domain, is a frequent solution. Existing methods such as mean-teacher, Cyclegan, and MCD will train models with the help of some easily accessible domains such as synthetic data. However, these methods are not guaranteed to operate well in real-world settings due to the domain shift. In this paper, we design a new unsupervised domain adaptation method named Multi-branch Adversarial Regressors (MarsDA) in hand pose estimation, where it could be better for feature migration. Specifically, we first generate pseudo-labels for unlabeled target domain data. Then, the new adversarial training loss between multiple regression branches we designed for hand pose estimation is introduced to narrow the domain gap. In this way, our model can reduce the noise of pseudo labels caused by the domain gap and improve the accuracy of pseudo labels. We evaluate our method on two publicly available real-world datasets, H3D and STB. Experimental results show that our method outperforms existing methods by a large margin. Jing Zhang 0037, Jianyu Yang 0002, Dacheng Tao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | A Multi-Level Network for Human Pose EstimationabstractAlthough multi-person human pose estimation has made great progress in recent years, the challenges such as various scales of persons, occluded keypoints, and crowded backgrounds in complex scenes are still remained to be solved. In this paper, we propose a novel multi-level pose estimation network (MLPE) to learn multi-level features that can preserve both the strong semantic clues and spatial resolution for keypoint prediction and location. More specifically, a multi-level prediction network with a feature enhancement strategy is first proposed to learn multi-level features to achieve a good trade-off between the global context information and spatial resolution. We then build a high-resolution fine network to restore high spatial resolution information based on transposed convolutions to accurately locate the keypoints. We have conducted extensive experiments on the challenging MS COCO dataset, which has proved the effectiveness of our proposed method. Code†and the experimental results are publicly online available for further research. Zhanpeng Shao, Youfu Li 0001, Jianyu Yang 0002, Xiaolong Zhou 0001 |
ICRA | 4 |
| 2021 | Learning discriminative motion feature for enhancing multi-modal action recognition
Jianyu Yang 0002, Zhanpeng Shao, Chunping Liu |
J. Vis. Commun. Image Represent. | 1 |
| 2021 | A multi-scale descriptor for real time RGB-D hand gesture recognition
Jianyu Yang 0002 |
Pattern Recognit. Lett. | 2 |
| 2021 | Hierarchical Soft Quantization for Skeleton-Based Human Action RecognitionabstractIn daily life, human beings rely on hands and body parts to complete particular actions cooperatively. These selected body parts and their cooperative relationships are essential cues to distinguish these actions. However, most existing action recognition methods, which try to model the body appearance or spatial relations in skeleton sequences, often ignore the essential cooperation relationship among joints. Differently, in this paper, we propose a spatio-temporal hierarchical soft quantization method to extract the congenerous motion features, which reflect the cooperation relations among joints and body parts. Specifically, we design a hierarchical network with multiple soft quantization layers to extract congenerous features. The hierarchical network not only models the spatial hierarchy of skeleton structure for joint, part, and body, but also extracts the temporal hierarchy with sliding windows for frame, fragment, and sequence. Moreover, the features in each layer are visually explainable, which reflect the cooperation among body parts. The trainable parameters in the network are also significantly reduced, which reduces computational cost. Extensive experiments conducted on four benchmarks demonstrate that our method can provide competitive results compared with state-of-the-arts. The visualized congenerous features also validate that our approach can effectively perceive the essential cooperation relations. Jianyu Yang 0002, Wu Liu 0005, Junsong Yuan 0001, Tao Mei 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | Selective Complementary Features For Multi-Person Pose EstimationabstractMulti-person pose estimation is a fundamental yet challenging research topic for many computer vision applications. It is difficult to achieve accurate localization results due to occlusion and complex background. In this paper, we propose a novel multi-person pose estimation approach with information complement and attention refinement residual module. To recover occlusion, the complementary features with multi-scale semantics information are extracted by our proposed Information Complement Module (ICM). To effectively discover the channel relationship and selectively highlight task-related regions in the feature maps, we design an Attention Refinement Residual Bottleneck (ARRB) module, which is an extension of residual unit with attention mechanism. We conduct ablation studies to investigate the efficacy of our method and compare it with the state-of-the-art methods on the COCO keypoint benchmark. Experimental results demonstrate that the selective complementary features are effective for multi-person pose estimation. Buwei Li, Yi Ji 0001, Jianyu Yang 0002, Chunping Liu |
ICIP | 4 |
| 2020 | Video Summarization with a Dual Attention Capsule NetworkabstractIn this paper, we address the problem of video summarization, which aims at selecting a subset of video frames as a summary to represent the original video contents compactly and completely. We propose a simple but effective supervised approach with a dual attention capsule network towards this end. Unlike existing LSTM based methods, it pays attention to short- and long-term dependencies among video frames through an elaborate dual self-attention architecture, which can handle longer-term dependencies and admit parallel computing. To reconcile the outputs of dual self-attention, we rely on a two-stream capsule network to learn the underlying frame selection criteria. Experiments on real-world datasets show the advantages of the proposed approach compared with state-of-the-art methods. Hongxing Wang 0001, Jianyu Yang 0002 |
ICPR | 3 |
| 2019 | Spatio-Temporal Multi-scale Soft Quantization Learning for Skeleton-Based Human Action RecognitionabstractEffective feature representation is important for action recognition. In this paper, a novel soft quantization learning method is proposed to represent visual features for action recognition. Specifically, we propose a dual multi-scale soft-quantization network, which is a trainable quantizer using RBF neurons. The RBF layer includes dual multi-scale structure, namely a three-level hierarchical skeleton structure in space, and a temporal-pyramid based multi-scale time structure. Different spatial levels in the RBF layer have respective RBF neurons for hierarchical spatial information, while the temporal scales share them to reduce the number of parameters in the network. An accumulation layer following the RBF layer summarizes the RBF output as a histogram representation for classification task. The proposed method is end-to-end differentiable that can be trained using regular back-propagation. The conducted experiments on benchmark datasets verify that the proposed method outperforms state-of-the-art methods. Jianyu Yang 0002, Junsong Yuan 0001 |
ICME | 1 |
| 2019 | Intra-Modality Feature Interaction Using Self-attention for Visual Question Answering
Huan Shao 0001, Yi Ji 0001, Jianyu Yang 0002, Chunping Liu |
ICONIP (5) | 4 |
| 2019 | Shape Description and Retrieval in a Fused Scale Space
Baojiang Zhong, Jianyu Yang 0002 |
ICONIP (2) | 3 |
| 2018 | Scene Graph Generation Based on Node-Relation Context Module
Chunping Liu, Yi Ji 0001, Jianyu Yang 0002 |
ICONIP (2) | 5 |
| 2018 | A Hierarchical Model for Action Recognition Based on Body PartsabstractAs increasing attention is paid on human action recognition from skeleton data, this paper focuses on such tasks by proposing a hierarchical model to discover the structure information of body-parts involved in human actions. Considering human actions as simultaneous motions of different body-parts of the human skeleton, we propose a hierarchical model to simultaneously apply discriminative body-parts selection at a same scale and group coupling of bundles of body-parts at different scales, while we decompose the human skeleton into a hierarchy of body-parts of varying scales. To represent such hierarchy of body-parts, we accordingly build a hierarchical RRV (Rotation and Relative Velocity) descriptors. The hierarchical representations encoded by Fisher vectors of the hierarchical RRV descriptors are properly formulated into the hierarchical model via the proposed hierarchical mixed norm, to apply sparse selection of body-parts and regularize the structure of such hierarchy of body-parts. The extensive evaluations on three challenging datasets demonstrate the effectiveness of our proposed approach, which achieves superior performance compared to state-of-the-art results on different sizes of datasets, showing it is more widely applicable than existing approaches. Zhanpeng Shao, Youfu Li 0001, Yao Guo 0002, Jianyu Yang 0002, Zhenhua Wang 0003 |
ICRA | 4 |
| 2017 | Real time hand gesture recognition via finger-emphasized multi-scale descriptionabstractThe development of depth cameras, e.g., the Kinect sensor, provides new opportunities for human computer interaction (HCI). Although the Kinect sensor has been extensively applied for human tracking, human action recognition and hand gesture recognition, real time hand gesture recognition is still a challenging problem. In this paper, we propose a new real time hand gesture recognition method. To represent the noisy and articulated hand shape segmented from the Kinect images, a finger emphasized multi-scale descriptor is proposed. To fully utilize hand shape features, this descriptor incorporates three types of parameters of multiple scales, which emphasize the finger features. Hand gesture recognition is then achieved with both DTW algorithm and BP neural network. Extensive experimental results and the comparison with state-of-the-art methods demonstrate that our method is accurate (a 100% accuracy on a challenging hand gesture dataset), efficient (average 0.941ms per frame), and robust to noise, articulations and rigid transformations. Jianyu Yang 0002, Junsong Yuan 0001 |
ICME | 1 |
| 2016 | Invariant multi-scale shape descriptor for object matching and recognitionabstractWe propose a novel multi-scale shape descriptor for shape matching and object recognition. The descriptor includes three types of invariants in multiple scales to capture discriminative local and semi-global shape features and the dynamic programming algorithm is employed for shape matching. The experimental results verify that our proposed shape feature is invariant to translation, rotation, scaling, and can well tolerate partial occlusion, articulated variation and intra-class variations. The shape matching and retrieval results on benchmark datasets validate the effectiveness of our method. This method is also applied to real time hand gesture recognition and achieves competitive result compared with state of the arts. Haoran Xu 0002, Jianyu Yang 0002, Junsong Yuan 0001 |
ICIP | 2 |
| 2016 | Invariant multi-scale descriptor for shape representation, matching and retrieval
Jianyu Yang 0002, Hongxing Wang 0001, Junsong Yuan 0001, Youfu Li 0001, Jianyang Liu |
Comput. Vis. Image Underst. | 1 |
| 2016 | Average consensus of multi-agent systems with self-triggered controllers
Jianyu Yang 0002 |
Neurocomputing | 2 |
| 2016 | Metric learning based object recognition and retrieval
Jianyu Yang 0002, Haoran Xu 0002 |
Neurocomputing | 1 |
| 2016 | Parsing 3D motion trajectory for gesture recognition
Jianyu Yang 0002, Junsong Yuan 0001, Youfu Li 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2015 | Flexible Trajectory Indexing for 3D Motion RecognitionabstractMotion trajectory analysis is important for human motion recognition and human computer interaction. In this paper, we propose a flexible 3D trajectory indexing method for complex 3D motion recognition. Based on both point level and primitive-level descriptors, trajectories are represented in the sub-primitive level, the level between the point level and primitive level. Primitives are flexibly segmented into sub-primitives in various scales, and the sub-primitives retain more detailed information than primitives. The detailed level of sub-primitives can be adjusted by controlling segmentation scales according to motion complexities. The proposed approach is suitable for spatial motion trajectory, which is view-invariant in 3D space. A cluster model is also proposed to represent motion classes and motion recognition performed based on maximum a posteriori (MAP) criterion. The experiments on benchmark datasets validate the effectiveness of the proposed approach. Jianyu Yang 0002, Junsong Yuan 0001, Youfu Li 0001 |
WACV | 1 |
| 2015 | Shape matching and object recognition using common base triangle areaabstractShape matching has always been a key issue in the field of computer vision. To obtain high recognition accuracy with low time complexity and to reduce the influence of contour deformation due to noise in shape matching, a novel shape matching method based on common base triangle area (CBTA) is proposed. First, a CBTA descriptor of each contour point is defined based on the area functions of the triangles formed by its two neighbour points and other contour points. Then, the descriptor is locally smoothed to keep it more compact and robust to noise. Secondly, a match cost matrix is obtained by computing the CBTA descriptors of all the contour points on two shapes. Finally, the similarity between the two shapes is measured on the basis of the match cost matrix by a dynamic programming algorithm. The experimental results on MPEG‐7, Kimia and an articulation shape database indicate that this method is robust to contour deformation, and both the computational efficiency and the retrieval rate are essentially improved. Dameng Hu, Weiguo Huang, Jianyu Yang 0002, Zhongkui Zhu |
IET Comput. Vis. | 3 |