EDBT 2026 Demo / reviewers in the wild / expert
Ke Xu 0003
dblp:181/2626-3
· DBLP profile ↗
38ranked-venue papers
6as first author
29since 2021 · last 2026
0000-0001-8771-9402ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Security and privacy · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When Does Adversarial Training Hurt Adversarial Robustness in Phase Translation of Training Data?
Ke Xu 0003, Xinghao Jiang, Tanfeng Sun, Zeyu Zhao 0006 |
ICIC (15) | 2 |
| 2026 | A Collaborative Adversarial Purification Framework with Perturbation-Insensitive Semantics-Guidance
Kaifeng Chen, Peisong He, Haoliang Li, Ke Xu 0003 |
ISCAS | 5 |
| 2026 | INN-RAE: Reversible adversarial examples based on invertible neural networks for facial protection
Zeyu Zhao 0006, Ke Xu 0003, Laijin Meng, Tanfeng Sun, Xinghao Jiang |
Expert Syst. Appl. | 2 |
| 2026 | Normalization-consistent data curation for generalizable deepfake detection
Shijie Hou, Xinghao Jiang, Ke Xu 0003, Qiang Xu 0007, Laijin Meng, Tanfeng Sun |
Neurocomputing | 3 |
| 2026 | PrePurify: Pre-trained knowledge-guided data purification for generalizable face forgery detection
Shijie Hou, Xinghao Jiang, Qiang Xu 0007, Ke Xu 0003, Laijin Meng |
Pattern Recognit. | 4 |
| 2026 | Adaptive Learning With Augmentation Robustness Validation: Toward Generalizable Face Forgery DetectionabstractThe rapid advancement of facial manipulation technologies demands detection systems that can generalize to novel forgery techniques. We identify that Standard Learning, reliant on static datasets and uniform sampling, detrimentally biases models towards specific patterns tied to individual generation techniques, hindering their ability to learn general features. To overcome this, we introduce Adaptive Learning (AL) for face forgery detection, a cyclical framework that simultaneously refines both the detector model and the training data through dynamic sample selection and model optimization. AL’s efficacy hinges on identifying samples rich in generalizable forgery clues. Thus, we propose Augmentation Robustness Validation (ARV) as AL’s core purification engine. ARV exploits the stability of predictions across diverse semantic-preserving augmentations as a reliable proxy for general feature presence: samples that exhibit invariant predictions inherently contain robust manipulation traces. Integrating ARV with AL yields Adaptive Learning with Augmentation Robustness Validation (ALarv). ALarv strategically prioritizes stability-verified samples during iterative training cycles, progressively enhancing the model’s focus on transferable forensic features. Inspired by the architectural advantages of ConvNeXt, we incorporate it into ALarv, forming an effective method, ALarv-ConvNeXt. Extensive experiments demonstrate ALarv-ConvNeXt’s superior generalization performance, including emerging diffusion-based synthetic faces. Shijie Hou, Xinghao Jiang, Ke Xu 0003, Qiang Xu 0007, Tanfeng Sun, Laijin Meng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | PSO-Based Closed Box Adversarial Patch Attack Against Face Recognition
Haotian Ma 0001, Xinghao Jiang, Ke Xu 0003, Tanfeng Sun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | ASGA: Attention-Based Sparse Global Attack to Video Action Recognition
Zeyu Zhao 0006, Ke Xu 0003, Tanfeng Sun, Xinghao Jiang |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2026 | SRAP: Robust and Transferable Self-Reversible Adversarial Patch for Image Privacy ProtectionabstractReversible adversarial examples offer adequate protection against malicious deep model identification and analysis. However, current methods still face challenges in terms of transferability and robustness, limiting their practical applicability. We introduce a novel technique for generating reversible adversarial examples utilizing Self-Reversible Adversarial Patch (SRAP) to address this. This approach significantly enhances the transferability and robustness of reversible adversarial examples against standard image processing techniques and adversarial defense methods. Specifically, we present a method for crafting adversarial patches that are small, non-overlapping, and adaptively integrated into specific regions. These adversarial patches are seamlessly combined with a reversible data-hiding technique that relies on prediction error expansion, resulting in adversarial examples with superior robustness and transferability. Experimental results indicate that our method achieves a remarkable transferability rate of up to 90% or higher between different models. Additionally, it exhibits strong robustness against image processing methods and adversarial defense strategies. Furthermore, our adversarial examples demonstrate an impressive attack success rate of 88% on commercial APIs, highlighting the effectiveness and practicality of our approach. Zeyu Zhao 0006, Ke Xu 0003, Tanfeng Sun, Xinghao Jiang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Identity-Agnostic Incremental Learning Framework for Face Forgery Detection
Jiayi Deng, Shuai Tang 0001, Ke Xu 0003, Peisong He |
PRCV (6) | 3 |
| 2025 | MV-guided deformable convolution network for compressed video action recognition with P-frames
Yuting Mou, Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
Neurocomputing | 2 |
| 2025 | Feature-Aware Transferable Adversarial Attacks on Visual Object TrackingabstractVisual object tracking is susceptible to adversarial attacks, posing significant security concerns for numerous application systems. Previous attack methods focused on white-box and untargeted attacks against response map. However, obtaining the tracking model in real-world scenarios is challenging, and the resulting adversarial trajectories are often unrealistic, making the attacks easily detectable. This paper proposes a Feature-aware Transferable Adversarial Patch (FTAP) that induces any black-box trackers to follow controllable and smooth trajectories. Tracker Following Assurance module is designed to manipulate bounding boxes to be valid and tightly align with the fake target. The movement of the tracker can be precisely controlled, resulting in adversarial trajectories stable and closely resemble natural trajectories, thereby reducing the risk of detection. The adversarial perturbation is generated solely from the initial template and applied to each frame. Consequently, the well-optimized generator can output universal adversarial patch capable of attacking any video without requiring additional computations. The intermediate layer features are corrupted to make the characteristics of the fake target closer to those of ground truth. Experimental results demonstrate that the proposed FTAP achieves state-of-the-art black-box attack performance and transferability across various tracker architectures. Mengdi Dong, Ke Xu 0003, Xinghao Jiang, Zeyu Zhao 0006, Tanfeng Sun |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | PLOVAD: Prompting Vision-Language Models for Open Vocabulary Video Anomaly DetectionabstractVideo anomaly detection (VAD) confronts significant challenges arising from data scarcity in real-world open scenarios, encompassing sparse annotations, labeling costs, and limitations on closed-set class definitions, particularly when scene diversity surpasses available training data. Although current weakly-supervised VAD methods offer partial alleviation, their inherent confinement to closed-set paradigms renders them inadequate in open-world contexts. Therefore, this paper explores open vocabulary video anomaly detection (OVVAD), leveraging abundant vision-related language data to detect and categorize both seen and unseen anomalies. To this end, we propose a robust framework, PLOVAD, designed to prompt tuning large-scale pretrained image-based vision-language models (I-VLMs) for the OVVAD task. PLOVAD consists of two main modules: the Prompting Module, featuring a learnable prompt to capture domain-specific knowledge and an anomaly-specific prompt crafted by a large language model (LLM) to capture semantic nuances and enhance generalization; and the Temporal Module, which integrates temporal information using graph attention network (GAT) stacking atop frame-wise visual features to address the transition from static images to videos. Extensive experiments on four benchmarks demonstrate the superior detection and categorization performance of our approach in the OVVAD task without bringing excessive parameters. Chenting Xu, Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | HEVC Video Adversarial Samples Detection via Joint Features of Compression and Pixel DomainsabstractDeep learning models are currently under significant threat from adversarial attacks, while adversarial detection represents an effective means of countering such assaults. However, existing adversarial detection techniques are deficient in localizing video adversarial frames, leading to poor performance on sparse video adversarial attacks. This paper presents an approach for detecting adversarial perturbations in videos based on fusion features derived from the video compression and RGB domain. Our research begins by examining how the introduction of extensive non-natural noise during video adversarial attacks severely disrupts the spatial structure of individual frames and the motion information between frames. This disruption culminates in unnatural variations in the Coding Tree Units (CTU) partitioning during the HEVC video encoding process. Then meticulously mapping the positions and partitioning information of coding units (CU), predictive units (PU), and transformation units (TU) onto specific values and sizes, constituting the video’s Compression Domain Units (CDU) features. Finally, a dual-path network utilizing both the video’s CDU features and the decoded frames RGB features is employed for detecting video adversarial samples. Extensive experiments are conducted to verify the performance. The results show that the proposed scheme outperforms or rivals the state-of-the-art methods in video adversarial detection. Zeyu Zhao 0006, Yueneng Wang, Ke Xu 0003, Tanfeng Sun, Xinghao Jiang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Preemptive Defense Algorithm Based on Generalizable Black-Box Feedback Regulation Strategy Against Face-Swapping Deepfake ModelsabstractIn the previous efforts to counteract Deepfake, detection methods were most adopted, but they could only function after-effect and could not undo the harm. Preemptive defense has recently gained attention as an alternative, but such defense works have either limited their scenario to facial-reenactment Deepfake models or only targeted specific face-swapping Deepfake model. Motivated to fill this gap, we start by establishing the Deepfake scenario modeling and finding the scenario difference among categories, then move on to the face-swapping scenario setting overlooked by previous works. Based on this scenario, we first propose a novel Black-Box Penetrating Defense Process that enables defense against face-swapping models without prior model knowledge. Then we propose a novel Double-Blind Feedback Regulation Strategy to solve the reality problem of avoiding alarming distortions after defense that had previously been ignored, which helps conduct valid preemptive defense against face-swapping Deepfake models in reality. Experimental results in comparison with state-of-the-art defense methods are conducted against popular face-swapping Deepfake models, proving our proposed method valid under practical circumstances. Zhongjie Mi, Xinghao Jiang, Tanfeng Sun, Ke Xu 0003, Qiang Xu 0007 |
IEEE Trans. Multim. | 4 |
| 2025 | Detection of HEVC Double Compression Based on Deep Representations of In-Loop Filtering and CU Depth MapsabstractIn the field of HEVC (High Efficiency Video Coding) double compression detection, relocated I-frame (RI frame) detection and original GOP size estimation are two significant problems for video forensics. However, little research explores the interconnection between the two problems, and effective methods to resolve them are still lacking. In this paper, a novel feature model called In-loop Filtering and CU Depth Map (IFCDM) is proposed to accurately detect RI frames, and the intrinsic correlation between RI frames and GOP structure is explored, which can be used for original GOP size estimation. Theoretical and statistical analysis of HEVC recompression process is first carried out. Then, sub-features of HEVC in-loop filtering modes and CU partition depth are extracted, and transformed into grey-scale maps to construct IFCDM. A neural network, consisting of tiny Vision Transformer and LSTM, is trained to learn spatial and temporal representations of input features, and further derive the RI frame detection results. Finally, an adaptive periodic analysis algorithm is designed, to integrate the RI frame detection results and estimate the original GOP size of recompressed videos. Experiments show that our method can outperform the existing state-of-the-art methods in both frame level and video level. Tanfeng Sun, Qiang Xu 0007, Ke Xu 0003, Xinghao Jiang |
IEEE Trans. Multim. | 4 |
| 2024 | Learning Spatio-Temporal Relations with Multi-Scale Integrated Perception for Video Anomaly DetectionabstractIn weakly supervised video anomaly detection, it has been verified that anomalies can be biased by background noise. Previous works attempted to focus on local regions to exclude irrelevant information. However, the abnormal events in different scenes vary in size, and current methods struggle to consider local events of different scales concurrently. To this end, we propose a multi-scale integrated perception (MSIP) learning approach to perceive abnormal regions of different scales simultaneously. In our method, a frame is partitioned into several groups of patches with varying scales, and a multi-scale patch spatial relation (MPSR) module is further proposed to model the inconsistencies among multi-scale patches. Specifically, we design a hierarchical graph convolution block in the MPSR module to improve the integration of patch features by implementing cross-scale feature learning. An existing clip temporal relation network is also introduced to enable spatio-temporal encoding in our model. Experiments show that our method achieves new state-of-the-art performance on the ShanghaiTech and competitive results on UCF-Crime benchmarks. Hongyu Ye, Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
ICASSP | 2 |
| 2024 | Low-Quality Deepfake Video Detection Model Targeting Compression-Degraded Spatiotemporal Inconsistencies
Zhongjie Mi, Xinghao Jiang, Tanfeng Sun, Ke Xu 0003, Qiang Xu 0007, Laijin Meng |
ICIC (9) | 4 |
| 2024 | Sparse Silhouette Jump: Adversarial Attack Targeted at Binary Image for Gait Privacy ProtectionabstractWith the widespread application of gait recognition technology, the issue of gait semantic security in videos has also attracted the attention of researchers. Its goal is to destroy the readability of data while preserving its semantic features. Thanks to the development of deep learning, some existing methods both domestically and internationally have made certain breakthroughs in recognition accuracy and visual effects. However, there is still significant room for improvement in balancing protection ability, visual effects, and computational complexity. This work is based on the ability of deep learning networks to extract gait identity features. In response to some problems in the current research field, we propose a gait privacy protection algorithm based on Sparse Silhouette Jump(SSJ), which draws on the idea of gradient descent in adversarial attacks and transfers adversarial noise to binary jumps to better adapt to binary graphs, while limiting the range of jumps from the perspective of spatial sparsity to balance the effectiveness and concealment of attacks. Experimental results have shown that our method achieves good effectiveness and concealment for various gait recognition models. Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
TrustCom | 2 |
| 2024 | Compressed Video Action Recognition Based on Neural Video CompressionabstractCompressed video action recognition based on traditional codecs, like MPEG-4, H265, etc., has achieved remarkable progress with comparable performance to raw video action recognition. With the development of Neural Video Compression (NVC), action recognition based on NVC should be paid attention to and explored. Firstly, the encoded stream of NVC represents the high-dimension features of the neural network, which allows the features to be utilized for downstream tasks with less additional processing or even directly. Secondly, the high-dimension feature can not be understood by humans, which means the privacy of the raw video frames can be preserved. In this paper, we propose a novel model for compressed video action recognition based on NVC to explore the potential of NVC for action recognition. By introducing spatial and temporal co-attention (ST-CA), the spatial information from the reference frame feature and the temporal information from the motion vector and the residual feature are combined and complemented effectively. The proposed model achieves competitive performance with the traditional compressed video action recognition methods and the raw video action recognition methods on the HMDB-51 and UCF-101 datasets. Besides, the proposed model preserves the privacy of the raw video frames and has much less computational complexity than the raw video action recognition methods. Yuting Mou, Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
TrustCom | 2 |
| 2024 | Compressed Video Action Recognition With Dual-Stream and Dual-Modal TransformerabstractCompressed video action recognition offers the advantage of reducing decoding and inference time compared to the RGB domain. However, the compressed domain poses unique challenges with different types of frames (I-frames and P-frames). I-frames consistent with RGB are rich in frame information, but the redundant information may interfere with the recognition task. There are two modalities in P-frames, residual (R) and motion vector (MV). Although with less information, they can reflect the motion cue. To address these challenges and leverage the independent information from different frames and modalities, we propose a novel approach called Dual-Stream and Dual-Modal Transformer (DSDMT). Our approach consists of two streams: 1) The short-span P-frames stream contains temporal information. We propose the Dual-Modal Attention Module (DAM) to mine different modal variability in P-frames and complement the orthogonal feature vector. Besides, considering the sparsity of P-frames, we extract action features with Frame-level Patch Embedding (FPE) to avoid redundant computation. 2) The long-span I-frames stream extracts the global context feature of the entire video, including content and scene information. By fusing the global video context and local key-frame features, our model represents the action feature in terms of fine-grained and coarse-grained. We evaluated our proposed DSDMT on three public benchmarks with different scales: HMDB-51, UCF-101, and Kinetics-400. Ours achieve better performance with fewer Flops and lower latency. Our analysis shows that the independence and complements of the I-frames and P-frames extracted from the compressed video stream play a crucial role in action recognition. Yuting Mou, Xinghao Jiang, Ke Xu 0003, Tanfeng Sun, Zepeng Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Feature Mixing and Disentangling for Occluded Person Re-IdentificationabstractOccluded person re-identification (Re-ID) has recently attracted lots of attention for its applicability in practical scenarios. However, previous pose-based methods always neglect the non-target pedestrian (NTP) problem. In contrast, we propose a feature mixing and disentangling method to train a robust network for occluded person Re-ID without extra data. Based on ViT, we design our network as follows: 1) A multi-target patch mixing (MPM) module is proposed to generate complex multi-target images with refined labels in the training stage. 2) We propose an identity-based patch realignment (IPR) module in the decoder layer to disentangle local features from the multi-target sample. In contrast to pose-guided methods, our approach overcomes the difficulties of NTP. More importantly, our approach does not bring additional computational costs in the training and testing phases. Experimental results show that our method effectively on occluded person Re-ID. For example, our method performs 3.3%/3.2% better than the baseline on Occluded-Duke in terms of mAP/rank-1 and outperforms the previous state-of-the-art. Zepeng Wang 0002, Ke Xu 0003, Yuting Mou, Xinghao Jiang |
ICME | 2 |
| 2023 | Deformable graph convolutional transformer for skeleton-based action recognition
Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
Appl. Intell. | 2 |
| 2023 | Transferable Black-Box Attack Against Face Recognition With Spatial Mutable Adversarial PatchabstractDeep Neural Networks (DNNs) are vulnerable to adversarial patch attacks, which raises security concerns for face recognition systems using DNNs. Previous attack methods focus on the perturbation texture and generate adversarial patches with fixed shapes at random or pre-designed locations, which causes poor adversarial transferability. This paper proposes a Spatial Mutable Adversarial Patch (SMAP) method to generate a dynamic mutable patch to be injected into the face. In the proposed SMAP, the texture, position and shape of the patch are optimized simultaneously and the patch generation pipeline is end-to-end differentiable. Specifically, a Patch Location Selection Scheme is designed to find the critical patch position with the most significant influence on the target identity by the step-based gradient search. By innovatively bridging the pre-defined mask and the dynamic update of the patch, the patch position and shape are changed based on the affine transformation and sampling mechanism in each iteration, which maintains the importance of the injected patch to the adversarial objective. To evaluate the vulnerability of face recognition models, we explore more threatening impersonation attacks under the black-box setting and design a strict evaluation metric that aligns with the real-world scenario. Extensive experiments show that the proposed SMAP improves attack performance across various face recognition models and datasets. Moreover, SMAP achieves better transferability on commercial face recognition systems than existing methods. Haotian Ma 0001, Ke Xu 0003, Xinghao Jiang, Zeyu Zhao 0006, Tanfeng Sun |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | A Transformer-Based Cloth-Irrelevant Patches Feature Extracting Method for Long-Term Cloth-Changing Person Re-identification
Zepeng Wang 0002, Xinghao Jiang, Ke Xu 0003, Tanfeng Sun |
CGI | 3 |
| 2022 | Research on video adversarial attack with long living cycleabstractIn recent years, the vulnerability of networks has attracted the attention of researchers. However, in these methods, the impact of video compression coding on the added adversarial perturbation, i.e., the robustness of the video adversarial example, is not considered. When an adversarial sample is just generated, its attack capability is the strongest. However, with multiple video encoding and video decoding in Internet transmission, the added adversarial disturbance will be continuously eliminated, eventually leading to the attack on the adversarial sample performance disappearing. We define this phenomenon as the decay of the lifetime of adversarial examples. We propose an adversarial attack method based on optimized integer space to resist this performance degradation. The robustness of anti-coding, the visual concealment, and the attack success rate are all considered during the attack process. In addition, we have also reduced the rounding loss caused by normalization in the deep neural network model process. The contributions of our methods are 1) We show the performance degradation caused by video compression coding on existing video adversarial attack methods, which seems an effective way for detecting of defending video adversarial examples. 2) A robust video adversarial attack method is proposed to resist video compression coding. The experiment shows that our method performs better on the robustness of anti-coding, visual concealment, and attack success rate. Zeyu Zhao 0006, Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
UAI | 2 |
| 2022 | Dual-domain graph convolutional networks for skeleton-based action recognition
Ke Xu 0003, Zhongjie Mi, Xinghao Jiang, Tanfeng Sun |
Mach. Learn. | 2 |
| 2022 | Gait Recognition Based on Local Graphical Skeleton Descriptor With Pairwise Similarity NetworkabstractGait recognition aims to identify a human through a walking sequence. It is a challenging task in computer vision since monocular camera loses most of the 3D information. Previous works described gait features with the contours of shape or the global geometrical characters of skeleton. So little work is researched on the local patterns of gait skeleton. In this paper, to resist the dress changes and speed changes, a Local Graphical Skeleton Descriptor (LGSD) is proposed to describe both the inner and intra local graphical patterns of a human gait skeleton. The gait features from the same or different identities are paired up and a Pairwise Similarity Network (PSN) is proposed to maximize the similarity of True matched pairs and minimize the similarity of False matched pairs. The contributions of our method are: 1) LGSD is proposed to describe human gait by computing four novel local geometrical patterns of skeleton sequences, which makes use of the intuitive cognition of gait based on the prior knowledge of mankind. 2) PSN is implemented by a two-stream CNN structure to build the gait model, which fused two popular gait recognition strategies. 3) The robustness of our method to dress changes and speed changes is proved on the public datasets. We have also achieved some state-of-the-art results on these datasets. The proposed method is examined on three public gait datasets which have RGB or infrared frames for evaluation: the CASIA-B dataset, the NLPR gait database, and the CASIA-C dataset. The performers in these datasets are walking under different views, speeds or dresses. The results are further compared with previous approaches to confirm the effectiveness and the advantages of our method. Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
IEEE Trans. Multim. | 1 |
| 2021 | Gait Identification Based on Human Skeleton with Pairwise Graph Convolutional NetworkabstractVision based gait identification is an important research content in the field of biometrics. Most existing gait identification methods extract representations from gait videos and identify a probe gait by ranking the similarities between the probe gait and all the gallery gaits. Since human body skeletons convey significant static and dynamic information of gait, in this paper, an end-to-end gait identification network named Pairwise Graph Convolutional Network (PGCN) is proposed to capture gait feature from skeletons. The skeleton sequences are first mapped into gait graphs and the PGCN is constructed to learn graph representations. The proposed method is examined on CASIA-B dataset on identical-view and cross-dress cases. The contributions of our work are: 1) The PGCN gait identification model is proposed to extract robust gait representation from skeleton sequences. 2) Different fusion structures are compared to explore the best fusion strategy for gait representation. 3) Using skeleton data, our work outperforms previous methods on identical-view and cross-dress cases. Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
ICME | 1 |
| 2020 | Action Recognition Scheme Based on Skeleton Representation With DS-LSTM NetworkabstractSkeleton-based human action recognition has been a popular research field during the past few years. With the help of cameras equipping deep sensors, such as the Kinect, human action can be represented by a sequence of human skeleton data. Inspired by the skeleton descriptors based on Lie group, a spatial-temporal skeleton transformation descriptor (ST-STD) is proposed in this paper. The ST-STD describes the relative transformations of skeletons, including the rotation and translation during movement. It gives a comprehensive view of the skeleton in both spatial and temporal domain for each frame. To capture the temporal connections in the skeleton sequence, a denoising sparse long short term memory (DS-LSTM) network is proposed in this paper. The DS-LSTM is designed to deal with two problems in action recognition. First, to decrease the intra-class diversity, the spatial-temporal auto-encoder (STAE) is proposed in this paper to generate representations with higher abstractness. The denoising constraint and the sparsity constraint are applied on both spatial and temporal domain to enhance the robustness and to reduce action misalignment. Second, to model the action sequence, a three-layer LSTM structure is trained with STAE representations for temporal modeling and classification. The experiments are carried out on four popular datasets. The results show that our approach performs better than several existing skeleton-based action recognition methods, which prove the effectiveness of our method. Xinghao Jiang, Ke Xu 0003, Tanfeng Sun |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Video Anomaly Detection and Localization Based on an Adaptive Intra-Frame Classification NetworkabstractVideo anomaly detection and localization is still a challenging task in the computer vision field. Previous methods took this task as an outlier detection problem, which computed the deviation between the test samples and the normal patterns. In this paper, an adaptive intra-frame classification network (AICN) is proposed to transform this task to a multi-class classification problem. The contributions of our method are as follows. AICN is an end-to-end network for anomaly detection and localization. By using the motion convolutional layers and the shape convolutional layers, spatial-temporal features are extracted without resizing or splitting the frames before forward propagation. AICN enhances the adaptiveness of model. By using the adaptive region pooling layer and the intra-frame classifier, AICN is adaptive to frames with different resolutions and is easier to be applied on other scenes. AICN evaluates the abnormality of frames based on the intra-frame classification results. The intra-frame classification strategy reserves more connection information of sub-regions and makes the model outperform previous methods. The proposed method is examined on four public datasets with different background complexities and resolutions: UCSD Ped1 dataset, UCSD Ped2 dataset, Avenue dataset and Subway dataset. The results are further compared with previous approaches to confirm the effectiveness and the advantage of our method. Ke Xu 0003, Tanfeng Sun, Xinghao Jiang |
IEEE Trans. Multim. | 1 |
| 2019 | Efficient Violence Detection Using 3D Convolutional Neural NetworksabstractAutomatically analyzing violent content in surveillance videos is of profound significance on many applications, ranging from Internet video filtration to public security protection. In this paper, we propose a deep learning model based on 3D convolutional neural networks, without using hand-crafted features or RNN architectures exclusively for encoding temporal information. The improved internal designs adopt compact but effective bottleneck units for learning motion patterns and leverage the DenseNet architecture to promote feature reusing and channel interaction, which is proved to be more capable of capturing spatiotemporal features and requires relatively fewer parameters. The performance of the proposed model is validated on three standard datasets in terms of recognition accuracy compared to other advanced approaches. Meanwhile, supplementary experiments are carried out to evaluate its effectiveness and efficiency. The final results demonstrate the advantages of the proposed model over the state-of-the-art methods in both recognition accuracy and computational efficiency. Xinghao Jiang, Tanfeng Sun, Ke Xu 0003 |
AVSS | 4 |
| 2019 | 3D Gait Recognition Based on a CNN-LSTM Network with the Fusion of SkeGEI and DA FeaturesabstractGait recognition is a promising technology in biometrics in video surveillance applications for its characteristics of non-contact and uniqueness. With the popularization of the Kinect sensor, human gait can be recognized based on the 3D skeletal information. For exploiting raw depth data captured by Kinect device effectively, a novel gait recognition approach based on Skeleton Gait Energy Image (SkeGEI) and Relative Distance and Angle (DA) features fusion is proposed. They are fused in backward to complement each other for gait recognition. In order to maintain as much gait information as possible, a CNN-LSTM network is designed to extract the temporal-spatial deep feature information from SkeGEI and DA features. The experiments evaluated on three datasets show that our approach performs superior to most gait recognition approaches with multi-directional and abnormal patterns. Xinghao Jiang, Tanfeng Sun, Ke Xu 0003 |
AVSS | 4 |
| 2019 | A Motion Vector-Based Steganographic Algorithm for HEVC with MTB Mapping Strategy
Mengyuan Guo, Tanfeng Sun, Xinghao Jiang, Ke Xu 0003 |
IWDW | 5 |
| 2019 | Gait Recognition with Clothing and Carrying Variations Based on GEI and CAPDS Features
Fengjia Yang, Xinghao Jiang, Tanfeng Sun, Ke Xu 0003 |
PRCV (2) | 4 |
| 2018 | Anomaly Detection Based on Stacked Sparse Coding With Intraframe Classification StrategyabstractAnomaly detection in videos is still a challenging task among the computer vision community. In this paper, an efficient anomaly detection method based on stacked sparse coding (SSC) with intraframe classification strategy is proposed. Each video is divided into blocks and the Foreground Interest Point (FIP) descriptor is proposed to describe the appearance and motion features for each block. The spatial-temporal features are then encoded with SSC. Specifically, the first stage of SSC encodes the spatial connections among blocks and the second stage of SSC encodes the temporal connections of all frame patches in each block. Finally, an intraframe classification strategy which uses the probabilistic outputs of SVM is proposed to evaluate the abnormality of each block. Contributions of this paper are listed as follows: 1) The FIP descriptor is proposed to describe the features of blocks, which reserves more spatial-temporal information. 2) The SSC encoding method encodes both the spatial and temporal connections of blocks, which makes the features more representative. 3) The intraframe classification strategy keeps the evaluation consistency among blocks and it helps to improve detection performance. The proposed method is examined on four public datasets with different background complexities and resolutions: UCSD Ped1 dataset, UCSD Ped2 dataset, Avenue dataset, and Subway dataset. The results are further compared with previous approaches to confirm the effectiveness and advantages of this method. Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
IEEE Trans. Multim. | 1 |
| 2017 | Two-Stream Dictionary Learning Architecture for Action RecognitionabstractIn this paper, a novel method based on the two-stream dictionary learning architecture for human action recognition is proposed. The architecture consists of interest patch (IP) detector and descriptor, two-stream dictionary models, and support vector machine (SVM) for classification. The novel IP detector combines a human detector and a contour detector to extract patches of interest on human contours. Then the IP descriptors are calculated in spatial stream and temporal stream separately. In each stream, a dictionary is trained for each action with the IP descriptors as an action model. In this way, measuring the similarity between an action sequence and an action model is transformed to reconstructing the IPs in this sequence with the model and computing the reconstruction error. For each action, an IP distribution histogram is constructed and the histogram is further used to train an SVM classifier in each stream. A score fusion method is applied to fuse the spatial and temporal SVM classification results to make a final decision. The proposed architecture is examined on four public data sets with different background complexities and camera motion conditions: Weizmann data set, KTH data set, Olympic sports data set, and HMDB51 data set. The results are further compared with state-of-the-art approaches in the experiment section to confirm the effectiveness of this architecture. Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Human activity recognition based on pose points selectionabstractA novel method for human action recognition is proposed in this paper. Traditional spatial-temporal interest point detectors are easily affected by hair, face, shadow, clothes texture or the shake of camera. Inspired by the use of points distribution information, we propose a point selection method to select representative points (denoted by the “pose points”), which use HOG human detector and contour detector to select the points on human pose edges. The pose points carry both local gradient information and global pose information. 3D-SIFT scale selection method and novel descriptors called body scale and motion intensity feature are also studied. The descriptors calculate the width scale of different levels of human body and count motion intensity of activity in five directions. The descriptors combine spatial location with the moving intensity together and are used for further classification with SVMs. Experiments have been conducted on benchmark datasets and show better performance than previous methods, which achieved 99.1% on Weizmann dataset and 95.8% on KTH dataset. Ke Xu 0003, Xinghao Jiang, Tanfeng Sun |
ICIP | 1 |