Siyuan Yang 0001

dblp:201/7699-1 · DBLP profile ↗
← Back
28ranked-venue papers
4as first author
26since 2021 · last 2026
0000-0003-4681-0431ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 16 · 4 first-author · 15 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond the Horizon: Decoupling Multi-View UAV Action Recognition via Partial Order Transfer
abstract
Action recognition using uncrewed aerial vehicles (UAVs) faces unique challenges due to substantial view variations along the vertical spatial axis. Unlike ground-based scenarios, UAVs capture actions from diverse altitudes, resulting in pronounced appearance discrepancies and reduced recognition robustness. To address this, we introduce a multi-view formulation tailored for UAV altitudes and empirically uncover a distinctive partial order among views, where recognition accuracy consistently declines as altitude increases. This key observation motivates the proposed Aero Partial Order Guided Network (Aerorder), which explicitly models and exploits the hierarchical structure of UAV views to enhance cross-altitude action recognition. Aerorder comprises three main components: (1) a View Partition (VP) module that groups views by altitude using the head-to-body ratio; (2) an Order-aware Feature Decoupling (OFD) module that disentangles action-relevant and view-specific representations under partial order guidance; and (3) an Action Partial Order Guide (APOG) that progressively transfers knowledge from easier (low-altitude) to harder (high-altitude) views. Extensive experiments on Drone-Action, MOD20, and UAV validate the superiority of Aerorder, achieving consistent improvements over state-of-the-art methods, up to 4.7% and 1.3% gains on Drone-Action and MOD20, respectively.
Wenxuan Liu 0008, Zhuo Zhou, Xuemei Jia, Siyuan Yang 0001, Wenxin Huang, Xian Zhong, Chia-Wen Lin
AAAI4
2026 AM40: Enhancing action recognition through matting-driven interaction analysis
Wenxuan Liu 0008, Kui Jiang, Siyuan Yang 0001, Chia-Wen Lin, Xian Zhong
Pattern Recognit.5
2026 Proactive Image Manipulation Detection and Tracing in Fake News
abstract
The pervasive spread of fake news, particularly through manipulated images, presents a consequential negative impact on society. To prevent fake news images from misleading the public, existing methods focus on verifying the authenticity of news images but ignore source traceability, leaving a gap in creating a complete forensic chain for reliable fake news detection. To simultaneously achieve the goals of authenticity verification and source tracing, we propose a proactive image tagging approach based on a design of Disentangled Invertible Neural Networks (DINN). It can simultaneously embed the dual-tags,i.e., authenticable tag and traceable tag, into each news image prior to publication, allowing for separate extraction for authenticity verification and source tracing. Within the proposed DINN, we design a parallel Feature Aware Projection Module (FAPM) to assist DINN in preserving essential tag information, thereby improving extraction accuracy. In addition, we introduce a Distance Metric-Guided Module (DMGM) that learns asymmetric one-class representations, enabling the dual-tags to exhibit different robustness performances under malicious manipulations. Extensive experiments on diverse datasets and unseen manipulations demonstrate that the proposed tagging approach achieves promising performances on both authenticity verification and source tracing for reliable fake news detection and outperforms the prior works.
Ruohan Meng, Siyuan Yang 0001, Zhili Zhou 0001, Kwok-Yan Lam, Zengwei Zheng, Alex Chichung Kot
IEEE Trans. Dependable Secur. Comput.2
2026 SpaceSeg: A High-Precision Intelligent Perception Segmentation Method for Multi-Spacecraft On-Orbit Targets
abstract
Accurate segmentation of multiple on-orbit spacecraft remains difficult in deep-space imagery because the scenes contain large uniform backgrounds, fine structural details, and limited labeled data. To address this problem, we propose SpaceSeg, a segmentation framework that adapts a vision foundation model to the spacecraft domain. The framework introduces a Multi-Scale Hierarchical Attention Refinement Decoder (MSHARD) to improve cross-scale feature decoding, a Spatial Domain Adaptation Transform (SDAT) training strategy to improve robustness to representative space-imaging disturbances, and a task-oriented objective that jointly optimizes segmentation accuracy and IoU-prediction quality. A lightweight connected-component-analysis module is also integrated into the pipeline for instance-aware target organization in multi-spacecraft scenes. We further construct SpaceES, a multi-scale on-orbit multi-spacecraft semantic segmentation dataset covering four space backgrounds and 17 spacecraft types. On SpaceES, SpaceSeg achieves 89.87% mIoU and 99.98% mAcc, setting a new state of the art among all evaluated baselines, surpassing the strongest competing method by 1.38 percentage points in mIoU with 59.6% fewer parameters, and exceeding the vanilla SAM2 baseline by 5.71 percentage points. Hardware-in-the-loop simulation and real satellite-to-satellite imagery experiments further support the practical relevance of the proposed method. Dataset and code are publicly available at https://github.com/Akibaru/SpaceSeg.
Pengyu Guo, Siyuan Yang 0001, Zeqing Jiang, Qinglei Hu, Dongyu Li
IEEE Trans. Image Process.3
2025 Adaptive Decision Boundary for Few-Shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) aims to continuously learn new classes from a limited set of training samples without forgetting knowledge of previously learned classes. Conventional FSCIL methods typically build a robust feature extractor during the base training session with abundant training samples and subsequently freeze this extractor, only fine-tuning the classifier in subsequent incremental phases. However, current strategies primarily focus on preventing catastrophic forgetting, considering only the relationship between novel and base classes, without paying attention to the specific decision spaces of each class. To address this challenge, we propose a plug-and-play Adaptive Decision Boundary Strategy (ADBS), which is compatible with most FSCIL methods. Specifically, we assign a specific decision boundary to each class and adaptively adjust these boundaries during training to optimally refine the decision spaces for the classes in each session. Furthermore, to amplify the distinctiveness between classes, we employ a novel inter-class constraint loss that optimizes the decision boundaries and prototypes for each class. Extensive experiments on three benchmarks, namely CIFAR100, miniImageNet, and CUB200, demonstrate that incorporating our ADBS method with existing FSCIL techniques significantly improves performance, achieving overall state-of-the-art results.
Linhao Li, Yongzhang Tan, Siyuan Yang 0001, Hao Cheng 0016, Yongfeng Dong
AAAI3
2025 Reconciling Stochastic and Deterministic Strategies for Zero-shot Image Restoration using Diffusion Model in Dual
abstract
Plug-and-play (PnP) methods offer an iterative strategy for solving image restoration (IR) problems in a zero-shot manner, using a learned discriminative denoiser as the implicit prior. More recently, a sampling-based variant of this approach, which utilizes a pre-trained generative diffusion model, has gained great popularity for solving IR problems through stochastic sampling. The IR results using PnP with a pre-trained diffusion model demonstrate distinct advantages compared to those using discriminative denoisers, i.e.,improved perceptual quality while sacrificing the data fidelity. The unsatisfactory results are due to the lack of integration of these strategies in the IR tasks. In this work, we propose a novel zero-shot IR scheme, dubbed Reconciling Diffusion Model in Dual (RDMD), which leverages only a single pre-trained diffusion model to construct two complementary regularizers. Specifically, the diffusion model in RDMD will iteratively perform deterministic denoising and stochastic sampling, aiming to achieve highfidelity image restoration with appealing perceptual quality. RDMD also allows users to customize the distortion-perception tradeoff with a single hyperparameter, enhancing the adaptability of the restoration process in different practical scenarios. Extensive experiments on several IR tasks demonstrate that our proposed method could achieve superior results compared to existing approaches on both the FFHQ and ImageNet datasets. Code is available at https://github.com/chongwang1024/rdmd.
Chong Wang 0011, Lanqing Guo, Zixuan Fu, Siyuan Yang 0001, Hao Cheng 0016, Alex Chichung Kot, Bihan Wen
CVPR4
2025 Vid-Group: Temporal Video Grounding Pretraining from Unlabeled Videos in the Wild
Peijun Bao, Chenqi Kong, Siyuan Yang 0001, Zihao Shao, Xinghao Jiang, Boon Poh Ng, Meng Hwa Er, Alex Chichung Kot
ICCV3
2025 MTL-UE: Learning to Learn Nothing for Multi-Task Learning
abstract
Most existing unlearnable strategies focus on preventing unauthorized users from training single-task learning (STL) models with personal data. Nevertheless, the paradigm has recently shifted towards multi-task data and multi-task learning (MTL), targeting generalist and foundation models that can handle multiple tasks simultaneously. Despite their growing importance, MTL data and models have been largely neglected while pursuing unlearnable strategies. This paper presents MTL-UE, the first unified framework for generating unlearnable examples for multi-task data and MTL models. Instead of optimizing perturbations for each sample, we design a generator-based structure that introduces label priors and class-wise feature embeddings which leads to much better attacking performance. In addition, MTL-UE incorporates intra-task and inter-task embedding regularization to increase inter-class separation and suppress intra-class variance which enhances the attack robustness greatly. Furthermore, MTL-UE is versatile with good supports for dense prediction tasks in MTL. It is also plug-and-play allowing integrating existing surrogate-dependent unlearnable methods with little adaptation. Extensive experiments show that MTL-UE achieves superior attacking performance consistently across 4 MTL datasets, 3 base UE methods, 5 model backbones, and 5 MTL task-weighting strategies. Code is available at https://github.com/yuyi-sd/MTL-UE.
Yi Yu 0011, Song Xia, Siyuan Yang 0001, Chenqi Kong, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot
ICML3
2025 HRHuman: Tuning-Free Higher-Resolution Human Image Generation via Template Knowledge
abstract
High-resolution human-centric image generation offers significant potential across various industries, such as entertainment, media, and fashion. Diffusion models for text-to-image generation have significantly improved the quality of human image synthesis. However, when scaling to higher resolutions (2K, 4K, and above), they often encounter issues such as object repetition and structural distortion, which appear especially unnatural in human images. To address these challenges, we propose HRHuman, a tuning-free framework for Higher-Resolution Human Image Generation. By leveraging an open-source large human vision model that incorporates rich template knowledge as prior, we first introduce the Prompt Discretization scheme to discretize user-input text prompts, mapping them to image elements and human body parts. Additionally, we implement a Template-guided Prompt Filtering mechanism to align these discretized prompts with regional image semantics, ensuring fine-grained prompt guidance. Extensive experiments demonstrate that HRHuman achieves state-of-the-art performance in human-centeric higher-resolution image generation, significantly addressing both issues of object repetition and structural distortion.
Ling Li 0012, Lanqing Guo, Siyuan Yang 0001, Yakun Ju, Weisi Lin, Alex Chichung Kot
ISCAS3
2025 Efficient Multi-Slide Visual-Language Feature Fusion for Placental Disease Classification
abstract
Accurate prediction of placental diseases via whole slide images (WSIs) is critical for preventing severe maternal and fetal complications. However, WSI analysis presents significant computational challenges due to the massive data volume. Existing WSI classification methods encounter critical limitations: (1) inadequate patch selection strategies that either compromise performance or fail to sufficiently reduce computational demands, and (2) the loss of global histological context resulting from patch-level processing approaches. To address these challenges, we propose an Efficient multimodal framework for Patient-level placental disease Diagnosis, named EmmPD. Our approach introduces a two-stage patch selection module that combines parameter-free and learnable compression strategies, optimally balancing computational efficiency with critical feature preservation. Additionally, we develop a hybrid multimodal fusion module that leverages adaptive graph learning to enhance pathological feature representation and incorporates textual medical reports to enrich global contextual understanding. Extensive experiments conducted on both a self-constructed patient-level Placental dataset and two public datasets demonstrating that our method achieves state-of-the-art diagnostic performance. The code is available at https://github.com/ECNU-MultiDimLab/EmmPD.
Zixuan Gao, Siyuan Yang 0001, Shulin Peng, Xiang Tao, Yan Wang 0033, Qingli Li
ACM Multimedia4
2025 Mining Generalized Multi-timescale Inconsistency for Detecting Deepfake Videos
Yang Yu 0039, Siyuan Yang 0001, Yu Ni, Yao Zhao 0001, Alex Chichung Kot
Int. J. Comput. Vis.3
2024 STSP: Spatial-Temporal Subspace Projection for Video Class-Incremental Learning
Hao Cheng 0016, Siyuan Yang 0001, Chong Wang 0011, Joey Tianyi Zhou, Alex Chichung Kot, Bihan Wen
ECCV (28)2
2024 Towards Physical World Backdoor Attacks Against Skeleton Action Recognition
Qichen Zheng, Yi Yu 0011, Siyuan Yang 0001, Jun Liu 0036, Kwok-Yan Lam, Alex Chichung Kot
ECCV (48)3
2024 Unlearnable Examples Detection via Iterative Filtering
Yi Yu 0011, Qichen Zheng, Siyuan Yang 0001, Wenhan Yang, Jun Liu 0036, Shijian Lu, Yap-Peng Tan, Kwok-Yan Lam, Alex Chichung Kot
ICANN (10)3
2024 One-Shot Action Recognition via Multi-Scale Spatial-Temporal Skeleton Matching
abstract
One-shot skeleton action recognition, which aims to learn a skeleton action recognition model with a single training sample, has attracted increasing interest due to the challenge of collecting and annotating large-scale skeleton action data. However, most existing studies match skeleton sequences by comparing their feature vectors directly which neglects spatial structures and temporal orders of skeleton data. This paper presents a novel one-shot skeleton action recognition technique that handles skeleton action recognition via multi-scale spatial-temporal feature matching. We represent skeleton data at multiple spatial and temporal scales and achieve optimal feature matching from two perspectives. The first is multi-scale matching which captures the scale-wise semantic relevance of skeleton data at multiple spatial and temporal scales simultaneously. The second is cross-scale matching which handles different motion magnitudes and speeds by capturing sample-wise relevance across multiple scales. Extensive experiments over three large-scale datasets (NTU RGB+D, NTU RGB+D 120, and PKU-MMD) show that our method achieves superior one-shot skeleton action recognition, and outperforms SOTA consistently by large margins.
Siyuan Yang 0001, Jun Liu 0036, Shijian Lu, Meng Hwa Er, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Self-Supervised 3D Action Representation Learning With Skeleton Cloud Colorization
abstract
3D Skeleton-based human action recognition has attracted increasing attention in recent years. Most of the existing work focuses on supervised learning which requires a large number of labeled action sequences that are often expensive and time-consuming to annotate. In this paper, we address self-supervised 3D action representation learning for skeleton-based action recognition. We investigate self-supervised representation learning and design a novel skeleton cloud colorization technique that is capable of learning spatial and temporal skeleton representations from unlabeled skeleton sequence data. We represent a skeleton action sequence as a 3D skeleton cloud and colorize each point in the cloud according to its temporal and spatial orders in the original (unannotated) skeleton sequence. Leveraging the colorized skeleton point cloud, we design an auto-encoder framework that can learn spatial-temporal features from the artificial color labels of skeleton joints effectively. Specifically, we design a two-steam pretraining network that leverages fine-grained and coarse-grained colorization to learn multi-scale spatial-temporal features. In addition, we design a Masked Skeleton Cloud Repainting task that can pretrain the designed auto-encoder framework to learn informative representations. We evaluate our skeleton cloud colorization approach with linear classifiers trained under different configurations, including unsupervised, semi-supervised, fully-supervised, and transfer learning settings. Extensive experiments on NTU RGB+D, NTU RGB+D 120, PKU-MMD, NW-UCLA, and UWA3D datasets show that the proposed method outperforms existing unsupervised and semi-supervised 3D action recognition methods by large margins and achieves competitive performance in supervised 3D action recognition as well.
Siyuan Yang 0001, Jun Liu 0036, Shijian Lu, Meng Hwa Er, Yongjian Hu, Alex Chichung Kot
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Cross-Modal Contrastive Pre-Training for Few-Shot Skeleton Action Recognition
abstract
This paper proposes a novel approach for few-shot skeleton action recognition that comprises of two stages: cross-modal pre-training of a skeleton encoder, followed by fine-tuning of a cosine classifier on the support set. The pre-training and fine-tuning approach has been demonstrated to be more effective for handling few-shot tasks compared to utilizing more intricate meta-learning methods. However, its success relies on the availability of a large-scale training dataset, which yet is difficult to obtain. To address this challenge, we introduce a cross-modal pre-training framework based on Bootstrap Your Own Latent (BYOL), which considers skeleton sequences and their corresponding videos as augmented views of the same action in different modalities. By utilizing a simple regression loss, the framework is able to transfer robust and high-quality vision-language representations to the skeleton encoder. This allows the skeleton encoder to gain a comprehensive understanding of action sequences and benefit from the prior knowledge obtained from a vision-language pre-trained model. The representation transfer enhances the feature extraction capability of the skeleton encoder, compensating for the lack of large-scale skeleton datasets. Extensive experiments on the NTU RGB+D, NTU RGB+D 120, PKU-MMD, NW-UCLA, and MSR Action Pairs datasets demonstrate that our proposed approach achieves state-of-the-art performances for few-shot skeleton action recognition.
Siyuan Yang 0001, Xiaobo Lu, Jun Liu 0036
IEEE Trans. Circuits Syst. Video Technol.2
2024 PVASS-MDD: Predictive Visual-Audio Alignment Self-Supervision for Multimodal Deepfake Detection
abstract
Deepfake techniques can forge the visual or audio signals in the video, which leads to inconsistencies between visual and audio (VA) signals. Therefore, multimodal detection methods expose deepfake videos by extracting VA inconsistencies. Recently, deepfake technology has started VA collaborative forgery to obtain more realistic deepfake videos, which poses new challenges for extracting VA inconsistencies. Recent multimodal detection methods propose to first extract natural VA correspondences in real videos in a self-supervised manner, and then use the learned real correspondences as targets to guide the extraction of VA inconsistencies in the subsequent deepfake detection stage. However, the inherent VA relations are difficult to extract due to the modality gap, which leads to the limited auxiliary performance of the aforementioned self-supervised methods. In this paper, we propose Predictive Visual-audio Alignment Self-supervision for Multimodal Deepfake Detection (PVASS-MDD), which consists of PVASS auxiliary and MDD stages. In the PVASS auxiliary stage in real videos, we first devise a three-stream network to associate two augmented visual views with corresponding audio clues, leading to explore common VA correspondences based on cross-view learning. Secondly, we introduce a novel cross-modal predictive align module for eliminating VA gaps to provide inherent VA correspondences. In the MDD stage, we propose to the auxiliary loss to utilize the frozen PVASS network to align VA features of real videos, to better assist multimodal deepfake detector for capturing subtle VA inconsistencies. We conduct extensive experiments on existing widely used and latest multimodal deepfake datasets. Our method obtains a significant performance improvement compared to state-of-the-art methods.
Yang Yu 0039, Siyuan Yang 0001, Yao Zhao 0001, Alex Chichung Kot
IEEE Trans. Circuits Syst. Video Technol.4
2024 Semantic Deep Hiding for Robust Unlearnable Examples
abstract
Ensuring data privacy and protection has become paramount in the era of deep learning. Unlearnable examples are proposed to mislead the deep learning models and prevent data from unauthorized exploration by adding small perturbations to data. However, such perturbations (e.g., noise, texture, color change) predominantly impact low-level features, making them vulnerable to common countermeasures. In contrast, semantic images with intricate shapes have a wealth of high-level features, making them more resilient to countermeasures and potential for producing robust unlearnable examples. In this paper, we propose a Deep Hiding (DH) scheme that adaptively hides semantic images enriched with high-level features. We employ an Invertible Neural Network (INN) to invisibly integrate predefined images, inherently hiding them with deceptive perturbations. To enhance data unlearnability, we introduce a Latent Feature Concentration module, designed to work with the INN, regularizing the intra-class variance of these perturbations. To further boost the robustness of unlearnable examples, we design a Semantic Images Generation module that produces hidden semantic images. By utilizing similar semantic information, this module generates similar semantic images for samples within the same classes, thereby enlarging the inter-class distance and narrowing the intra-class distance. Extensive experiments on CIFAR-10, CIFAR-100, and an ImageNet subset, against 18 countermeasures, reveal that our proposed method exhibits outstanding robustness for unlearnable examples, demonstrating its efficacy in preventing unauthorized data exploitation.
Ruohan Meng, Chenyu Yi, Yi Yu 0011, Siyuan Yang 0001, Bingquan Shen, Alex Chichung Kot
IEEE Trans. Inf. Forensics Secur.4
2024 Progressive Channel-Shrinking Network
abstract
Currently, salience-based channel pruning makes continuous breakthroughs in network compression. In the realization, the salience mechanism is used as a metric of channel salience to guide pruning. Therefore, salience-based channel pruning can dynamically adjust the channel width at run-time, which provides a flexible pruning scheme. However, there are two problems emerging: a gating function is often needed to truncate the specific salience entries to zero, which destabilizes the forward propagation; dynamic architecture brings more cost for indexing in inference which bottlenecks the inference speed. In this article, we propose a Progressive Channel-Shrinking (PCS) method to compress the selected salience entries at run-time instead of roughly approximating them to zero. We also propose a Running Shrinking Policy to provide a testing-static pruning scheme that can reduce the memory access cost for filter indexing. We evaluate our method on ImageNet and CIFAR10 datasets over two prevalent networks: ResNet and VGG, and demonstrate that our PCS outperforms all baselines and achieves state-of-the-art in terms of compression-performance tradeoff. Moreover, we observe a significant and practical acceleration of inference. The code is available athttps://github.com/JianhongPan-VLG/Progressive.Channel-Shrinking.Network.
Jianhong Pan, Siyuan Yang 0001, Lin Geng Foo, Qiuhong Ke, Hossein Rahmani 0001, Zhipeng Fan 0001, Jun Liu 0036
IEEE Trans. Multim.2
2024 Narrowing Domain Gaps With Bridging Samples for Generalized Face Forgery Detection
abstract
Face forgery technology has developed rapidly, causing severe security issues in society. Recently, with the continuous emergence of forgery techniques and types, most forensics methods suffer from the generalization problem. In particular, it is difficult for existing generalized methods to detect fake faces with unseen fake types. The reason is that the distribution gaps among cross-forgery types are too large. In this article, we propose a novel generalized framework to narrow large gaps based on bridging cross-domain alignment to solve this problem. Specifically, our framework consists of three key steps: preventing, bridging and aligning distribution gaps. Firstly, in the feature mining stage, taking advantage of the ability of Instance Normalization (IN) to better tolerate domain gaps, we design Adaptive Batch and Instance Normalization (ABIN) to replace the commonly used BN to adaptively extract features to preliminarily prevent domain gaps. Secondly, we propose to generate bridging samples distributed among the inter-domains to fill large gaps based on progressive linear interpolation operation. Finally, with the help of bridging samples, the cross-domain alignment is performed to better narrow distribution gaps to refine data distribution, which helps to learn a more generalized framework. Extensive experiments show that our proposed framework achieves the state-of-the-art generalized performance.
Yang Yu 0039, Siyuan Yang 0001, Yao Zhao 0001, Alex Chichung Kot
IEEE Trans. Multim.3
2023 Frequency Guidance Matters in Few-Shot Learning
abstract
Few-shot classification aims to learn a discriminative feature representation to recognize unseen classes with few labeled support samples. While most few-shot learning methods focus on exploiting the spatial information of image samples, frequency representation has also been proven essential in classification tasks. In this paper, we investigate the effect of different frequency components on the few-shot learning tasks. To enhance the performance and generalizability of few-shot methods, we propose a novel Frequency-Guided Few-shot Learning framework (dubbed FGFL), which leverages the task-specific frequency components to adaptively mask the corresponding image information, with a novel multi-level metric learning strategy including a triplet loss among original, masked and unmasked image as well as a contrastive loss between masked and original support and query sets to exploit more discriminative information. Extensive experiments on four benchmarks under several few-shot scenarios, i.e., standard, cross-dataset, cross-domain, and coarse-to-fine annotated classification, are conducted. Both qualitative and quantitative results show that our proposed FGFL scheme can attend to the class-discriminative frequency components, thus integrating those information towards more effective and generalizable few-shot learning.
Hao Cheng 0016, Siyuan Yang 0001, Joey Tianyi Zhou, Lanqing Guo, Bihan Wen
ICCV2
2023 Temporal Coherent Test Time Optimization for Robust Video Classification
Chenyu Yi, Siyuan Yang 0001, Yufei Wang 0006, Haoliang Li, Yap-Peng Tan, Alex Chichung Kot
ICLR2
2023 MSVT: Multiple Spatiotemporal Views Transformer for DeepFake Video Detection
abstract
Recently, DeepFake videos have developed rapidly, causing new security issues in society. Due to the rough spatiotemporal view, existing video-based detection methods struggle to capture fine-grained spatiotemporal information, resulting in limited generalization ability. In addition, although the transformer has achieved great success in the past few years, the application of transformer on deepfake video detection still needs to be studied. To solve this problem, in this paper, we propose a novel Multiple Spatiotemporal Views Transformer (MSVT) with Local Spatiotemporal View (LSV) and Global Spatiotemporal View (GSV), to mine more detailed spatiotemporal information. Firstly, for establishing the LSV, different from existing works that sparsely sample a single frame to build the input sequence, we employ the local-consecutive temporal view to capture vital dynamic inconsistency. Furthermore, the extracted frame features within each group are fed to the temporal transformer followed by the feature fusion module, to generate group-level spatiotemporal features. Then, we further establish Global Spatiotemporal View (GSV) by feeding all the frame features within the whole video to the temporal transformer followed by the feature fusion module. Finally, we propose a novel global-local transformer (GLT) to effectively integrate these multi-level features for mining more subtle and comprehensive features. Extensive experiments on six large datasets demonstrate that our MSVT outperforms state-of-the-art detection methods.
Yang Yu 0039, Yao Zhao 0001, Siyuan Yang 0001, Fen Xia
IEEE Trans. Circuits Syst. Video Technol.4
2023 Augmented Multi-Scale Spatiotemporal Inconsistency Magnifier for Generalized DeepFake Detection
abstract
Recently, realistic DeepFake videos have raised severe security concerns in society. Existing video-based detection methods observe local spatial regions with the coarse temporal view, thus it is difficult to obtain subtle spatiotemporal information, resulting in limited generalization ability. In this paper, we propose a novel Augmented Multi-scale Spatiotemporal Inconsistency Magnifier (AMSIM) with a Global Inconsistency View (GIV) and a more meticulous Multi-timescale Local Inconsistency View (MLIV), focusing on mining comprehensive and more subtle spatiotemporal cues. Firstly, the GIV that includs the global spatial and long-term temporal views is established to ensure comprehensive spatiotemporal clues are captured. Then, the MLIV with the critical local spatial and multi-timescale local temporal views is designed for magnifying the indetectable spatiotemporal abnormality. Subsequently, GIV is utilized to guide MLIV to dynamically find local spatiotemporal anomalies that are highly relevant to the overall video. Finally, to further obtain a generalized framework, the adversarial data augmentation is specially designed to expand source domains and simulate unseen forgery domains. Extensive experiments on six large-scale datasets show that our AMSIM outperforms state-of-the-art detection methods and remains effective when applied to unseen forgery techniques and datasets.
Yang Yu 0039, Siyuan Yang 0001, Yao Zhao 0001, Alex Chichung Kot
IEEE Trans. Multim.4
2021 Skeleton Cloud Colorization for Unsupervised 3D Action Representation Learning
abstract
Skeleton-based human action recognition has attracted increasing attention in recent years. However, most of the existing works focus on supervised learning which requiring a large number of annotated action sequences that are often expensive to collect. We investigate unsupervised representation learning for skeleton action recognition, and design a novel skeleton cloud colorization technique that is capable of learning skeleton representations from unlabeled skeleton sequence data. Specifically, we represent a skeleton action sequence as a 3D skeleton cloud and colorize each point in the cloud according to its temporal and spatial orders in the original (unannotated) skeleton sequence. Leveraging the colorized skeleton point cloud, we design an auto-encoder framework that can learn spatial-temporal features from the artificial color labels of skeleton joints effectively. We evaluate our skeleton cloud colorization approach with action classifiers trained under different configurations, including unsupervised, semi-supervised and fully-supervised settings. Extensive experiments on NTU RGB+D and NW-UCLA datasets show that the proposed method outperforms existing unsupervised and semi-supervised 3D action recognition methods by large margins, and it achieves competitive performance in supervised 3D action recognition as well.
Siyuan Yang 0001, Jun Liu 0036, Shijian Lu, Meng Hwa Er, Alex Chichung Kot
ICCV1
2020 Collaborative Learning of Gesture Recognition and 3D Hand Pose Estimation with Multi-order Feature Analysis
Siyuan Yang 0001, Jun Liu 0036, Shijian Lu, Meng Hwa Er, Alex Chichung Kot
ECCV (3)1
2020 Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark Study
abstract
Existing enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions.
Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin
IEEE Trans. Image Process.18