Jinyu Tian 0001

dblp:14/1023-1 · DBLP profile ↗
← Back
41ranked-venue papers
4as first author
37since 2021 · last 2026
0000-0002-2449-5277ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 2 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Security and privacy · 4 · 4 since 2021Computer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Deferred Poisoning: Making the Model More Vulnerable via Hessian Singularization
abstract
Recent studies have shown that deep learning models are very vulnerable to poisoning attacks. Many defense methods have been proposed to address this issue. However, traditional poisoning attacks are not as threatening as commonly believed. This is because they often cause differences in how the model performs on the training set compared to the validation set. Such inconsistency can alert defenders that their data has been poisoned, allowing them to take the necessary defensive actions. In this paper, we introduce a more threatening type of poisoning attack called the Deferred Poisoning Attack. This new attack allows the model to function normally during the training and validation phases but makes it very sensitive to evasion attacks or even natural noise. We achieve this by ensuring the poisoned model's loss function has a similar value as a normally trained model at each input sample but with a large local curvature. A similar model loss ensures that there is no obvious inconsistency between the training and validation accuracy, demonstrating high stealthiness. On the other hand, the large curvature implies that a small perturbation may cause a significant increase in model loss, leading to substantial performance degradation, which reflects a worse robustness. We fulfill this purpose by making the model have singular Hessian information at the optimal point via our proposed Singularization Regularization term. We have conducted both theoretical and empirical analyses of the proposed method and validated its effectiveness through experiments on image classification tasks. Furthermore, we have confirmed the hazards of this form of poisoning attack under more general scenarios using natural noise, offering a new perspective for research in the field of security.
Yuhao He 0001, Jinyu Tian 0001, Xianwei Zheng, Li Dong 0006, Yuanman Li, Jiantao Zhou 0001
AAAI2
2026 OAD-Promoter: Enhancing Zero-Shot VQA Using Large Language Models with Object Attribute Description
abstract
Large Language Models (LLMs) have become a crucial tool in Visual Question Answering (VQA) for handling knowledge-intensive questions in few-shot or zero-shot scenarios. However, their reliance on massive training datasets often causes them to inherit language biases during the acquisition of knowledge. This limitation imposes two key constraints on existing methods: (1) LLM predictions become less reliable due to bias exploitation, and (2) despite strong knowledge reasoning capabilities, LLMs still struggle with out-of-distribution (OOD) generalization. To address these issues, we propose Object Attribute Description Promoter (OAD-Promoter), a novel approach for enhancing LLM-based VQA by mitigating language bias and improving domain-shift robustness. OAD-Promoter comprises three components: the Object-concentrated Example Generation (OEG) module, the Memory Knowledge Assistance (MKA) module, and the OAD Prompt. The OEG module generates global captions and object-concentrated samples, jointly enhancing visual information input to the LLM and mitigating bias through complementary global and regional visual cues. The MKA module assists the LLM in handling OOD samples by retrieving relevant knowledge from stored examples to support questions from unseen domains. Finally, the OAD Prompt integrates the outputs of the preceding modules to optimize LLM inference. Experiments demonstrate that OAD-Promoter significantly improves the performance of LLM-based VQA methods in few-shot or zero-shot settings, achieving new state-of-the-art results.
Quanxing Xu, Ling Zhou 0005, Feifei Zhang 0001, Rubing Huang, Jinyu Tian 0001
AAAI5
2026 Learning a Semantic Similarity Orthogonal Space for Model-Level AI -Generated Image Source Attribution
abstract
ABSTRACT The rapid evolution of diffusion models has established artificial intelligence generated image (AIGI) as a dominant paradigm in digital media, simultaneously escalating the risks of deepfakes and copyright infringement. While existing forensic methods focus primarily on distinguishing real from synthetic content, they fail to address the critical attribution challenge: identifying the specific model architecture responsible for a generated image. To bridge this gap, this paper proposes a novel AIGI source tracing framework capable of pinpointing source models in black‐box scenarios. Our approach is grounded in the hypothesis that images generated by the same model under identical semantic conditions exhibit consistent stylistic signatures and detail‐level artefacts. The framework operates through a ‘reconstruct‐and‐compare’ paradigm. First, we employ a CLIP‐based optimization algorithm to reconstruct the semantic prompt of the query image, enabling the generation of a reference sample from the candidate model. Second, to accurately measure the provenance similarity between the query and the reference, we introduce a Semantic Comparison Network based on Orthogonal Extension. This network utilizes a feature purification mechanism to decouple shared semantic content from model‐specific visual attributes. Furthermore, it employs a multi‐stage orthogonal training strategy to extract mutually independent feature subspaces, ensuring a comprehensive capture of diverse stylistic dimensions. Experimental results demonstrate that this framework effectively overcomes the limitations of current classifier‐based attribution, offering a robust solution for digital forensics and accountable AI governance.
Yuchu Lin, Xinqi Jiang, Wei Wang 0077, Jinyu Tian 0001
Expert Syst. J. Knowl. Eng.4
2026 Short-term electricity load forecasting with multi-frequency reconstruction diffusion
Rubing Huang, Ling Zhou 0005, Dave Towey, Jinyu Tian 0001
Inf. Sci.5
2026 Robust graph neural networks via supervised block diagonal regularizer
Yulong Wang 0002, Huiwu Luo, Jinyu Tian 0001, Yuan Yan Tang
Pattern Recognit.7
2026 ETV-Attack: Efficient text-driven visual-variable adversarial attacks on visual question answering with pre-trained language models
Quanxing Xu, Ling Zhou 0005, Xian Zhong, Feifei Zhang 0001, Jinyu Tian 0001, Xiaohan Yu 0001, Rubing Huang
Pattern Recognit.5
2026 Refined generation-based framework for consistent and reliable visual question answering
Quanxing Xu, Ling Zhou 0005, Xian Zhong, Feifei Zhang 0001, Jinyu Tian 0001, Xiaohan Yu 0001, Rubing Huang
Pattern Recognit.5
2026 Reversible Unlearnable Examples: Toward the Copyright Protection in Deep Learning Era
abstract
Significant advancements in deep learning have been made possible by the utilization of large datasets, underscoring the critical importance of copyright protection. Adding meticulously designed perturbations to examples, making them unlearnable has become a crucial approach for safeguarding data copyright. Existing methods for creating unlearnable examples overlook the risk of data leakage, which can threaten data ownership. Thus, copyright protection in deep learning faces two main threats: illegal model training and malicious data leakage. We investigate that these two threats cannot be solved by straightforwardly combining existing availability attacks and watermarking techniques as their negative interaction effects. Therefore, in this paper, we propose a novel copyright protection mechanism for the aforementioned security concerns. Considering that the prevention of unauthorized model training requires powerful generalizability of unlearnable perturbations, we generate perturbations to induce the model to learn uncorrelated features of input images. It works by minimizing the mutual information of the input and output of the model. On the other hand, to eliminate the side impact of unlearnable perturbations on the watermark extraction, we design a dual extraction strategy by using two distinct watermark extractors. Extensive experiments on the image datasets ImageNet, CIFAR10, and Pets show that our proposed method could provide comprehensive copyright protection to images. The code is available at https://github.com/Yeah21/ReversibleUnlearnableExamples.
Binze Wang, Jinyu Tian 0001, Xingrun Wang, Xiaochen Yuan, Jianqing Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 DFFormer: Capturing Dynamic Frequency Features to Locate Image Manipulation Through Adaptive Frequency Transformer and Prototype Learning
abstract
The proliferation of modern image editing tools has raised concerns about image manipulation, particularly regarding the potential to mislead the public and compromise privacy and security. Consequently, detecting and localizing tampered regions has become a critical research challenge. Traditional methods struggle with subtle manipulations, such as splicing, copy-move, and removal, which are often more discernible in the frequency domain than in the spatial domain. Additionally, the size imbalance between the tampered and background regions further complicates the detection process. To address these challenges, we propose DFFormer, an end-to-end network that leverages frequency feature differences and a dynamic token strategy for precise manipulation localization. DFFormer combines the Conventional Neural Network (CNN) and Transformer in a hybrid architecture with three key modules: the Adaptive Frequency Transformer (AFT), the Prototype Learning Module (PLM), and the Cascaded Progressive Token Fusion Head (CPTF-Head). AFT integrates high- and low-frequency components into self-attention via the Parallel Adaptive Frequency Attention (PAFA) block, enhancing tampering feature representation while preserving fine details. PLM employs KNN-based density peak clustering (DPC-KNN) and weighted token aggregation to optimize dynamic token reduction. The CPTF-Head adopts a hierarchical coarse-to-fine strategy to integrate multiscale features, thereby improving localization accuracy and edge refinement. Experiments demonstrate that DFFormer outperforms state-of-the-art models across four benchmark datasets and one real-world dataset, exhibiting superior generalization and robustness. The source code is publicly available at https://github.com/XiangGD/DFFormer.git.
Kaiqi Zhao 0004, Zhenghong Yu, Xiaochen Yuan, Guoheng Huang, Jinyu Tian 0001, Jianqing Li 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Anti-Diffusion: Preventing Abuse of Modifications of Diffusion-Based Models
abstract
Although diffusion-based techniques have shown remarkable success in image generation and editing tasks, their abuse can lead to severe negative social impacts. Recently, some works have been proposed to provide defense against the abuse of diffusion-based methods. However, their protection may be limited in specific scenarios by manually defined prompts or the stable diffusion (SD) version. Furthermore, these methods solely focus on tuning methods, overlooking editing methods that could also pose a significant threat. In this work, we propose Anti-Diffusion, a privacy protection system designed for general diffusion-based methods, applicable to both tuning and editing techniques. To mitigate the limitations of manually defined prompts on defense performance, we introduce the prompt tuning (PT) strategy that enables precise expression of original images. To provide defense against both tuning and editing methods, we propose the semantic disturbance loss (SDL) to disrupt the semantic information of protected images. Given the limited research on the defense against editing methods, we develop a dataset named Defense-Edit to assess the defense performance of various methods. Experiments demonstrate that our Anti-Diffusion achieves superior defense performance across a wide range of diffusion-based techniques in different scenarios.
Liangbin Xie, Jiantao Zhou 0001, Xintao Wang 0002, Haiwei Wu, Jinyu Tian 0001
AAAI6
2025 Data-Free Universal Attack by Exploiting the Intrinsic Vulnerability of Deep Models
abstract
Deep neural networks (DNNs) are susceptible to Universal Adversarial Perturbations (UAPs), which are instance-agnostic perturbations that can deceive a target model across a wide range of samples. Unlike instance-specific adversarial examples, UAPs present a greater challenge as they must generalize across different samples and models. Generating UAPs typically requires access to numerous examples, which is a strong assumption in real-world tasks. In this paper, we propose a novel data-free method called Intrinsic UAP(IntriUAP), by exploiting the intrinsic vulnerabilities of deep models. We analyze a series of popular deep models composed of linear and nonlinear layers with a Lipschitz constant of 1, revealing that the vulnerability of these models is predominantly influenced by their linear components. Based on this observation, we leverage the ill-conditioned nature of the linear components by aligning the UAP with the right singular vectors corresponding to the maximum singular value of each linear layer. Remarkably, our method achieves highly competitive performance in attacking popular image classification deep models without using any image samples. We also evaluate the black-box attack performance of our method, showing that it matches the state-of-the-art baseline for data-free methods on models that conform to our theoretical framework. Beyond the data-free assumption, IntriUAP also operates under a weaker assumption, where the adversary only can access a few of the victim model's layers. Experiments demonstrate that the attack success rate decreases by only 4% when the adversary has access to just 50% of the linear layers in the victim model.
YangTian Yan, Jinyu Tian 0001
AAAI2
2025 Short-Term Electricity-Load Forecasting by deep learning: A comprehensive survey
Rubing Huang, Chenhui Cui, Dave Towey, Ling Zhou 0005, Jinyu Tian 0001
Eng. Appl. Artif. Intell.6
2025 RML++: Regroup Median Loss for Combating Label Noise
Fengpeng Li, Kemou Li, Bo Han 0003, Jinyu Tian 0001, Jiantao Zhou 0001
Int. J. Comput. Vis.5
2025 Security Within Security: Attack Detection Model With Defenses Against Attacks Capability for Zero-Trust Networks
abstract
Traditional traffic anomaly-based attack detection methods in Zero-trust Networks (ZTN) suffer from inherent security vulnerabilities, as they neglect considerations regarding their security defenses. Compromising the attack detection model itself can result in the breakdown of normal attack detection capabilities. Ensuring the security of the attack detection model during runtime presents a novel challenge. To address these shortcomings, we propose a novel attack detection model, termed Security within Security: Attack Detection Model with Defenses Against Attacks Capability for Zero-Trust Networks (SWS), aimed at enhancing the security of ZTN. SWS focuses on achieving attack detection in non-secure detection environments, to maintain its detection capability even when under attack. By employing a soft thresholding method, SWS adapts to the dynamic changes in network traffic, thus reducing the interference of attack signals. The incorporation of an attention mechanism enables SWS to concentrate on analyzing the most indicative traffic features of attack behavior. Additionally, we integrate Residual Networks (ResNet) and Bidirectional Long Short-Term Memory (BiLSTM) to enhance the robustness of identifying complex network attack behaviors. The effectiveness of the SWS is validated through ablation studies, model comparisons, experiments conducted over different training epochs, and experiments conducted on various components of the dataset. Experimental results demonstrate that compared to existing attack detection models, SWS achieves improvements in detection accuracy and recall rate by 13.4% and 10.6%, respectively, while reducing the False Positive Rate (FPR) by 16.9%.
Tingting Wang 0006, Kai Fang 0001, Jijing Cai, Jinyu Tian 0001, Hailin Feng, Jianqing Li 0001, Mohsen Guizani, Wei Wang 0077
IEEE J. Sel. Areas Commun.5
2025 Toward Robust Learning via Core Feature-Aware Adversarial Training
abstract
Deep neural networks (DNNs) are inherently vulnerable to adversarial examples (AEs), severely deteriorating model performance on various tasks. Adversarial training (AT) is one of the most effective approaches to enhance model robustness by incorporating AEs into the training process. Notwithstanding the efficacy of AT, recent studies have unveiled that adversarial perturbations on AEs predominantly impact core features—essential for accurate predictions—more than spurious features, which are incidentally aligned with training labels but irrelevant to the model’s classification. This unequal impact induces the models trained with AT to excessively rely on spurious features, resulting in a pronouncedfeature shiftthat compromises robustness and generalization against AEs at inference. In this work, we introduce a novelCore Feature-aware Adversarial Training(COFAT) framework to cope with these challenges. COFAT employscore feature extractionto dynamically generatecore partnersby selectively retaining benign sample regions on feature maps with high-weight while masking low-weight ones, thereby ensuring the model focuses on core features. Furthermore,contrastive feature alignmentis proposed to reduce intra-class feature distances and increase inter-class separability by maintaining a center bank of class feature representations, thus mitigating reliance on spurious features. Compared to state-of-the-art AT methods, COFAT demonstrates superior performance against diverse adversarial attacks. Remarkably, COFAT improves the robustness of ResNet-18 against AutoAttack on CIFAR-10, SVHN, CIFAR-100, and Tiny ImageNet by approximately 2.14%, 3.20%, 1.69%, and 1.86%, respectively, embodying significant advancements in AT. Our code is publicized at https://github.com/Feng-peng-Li/CoFAT.
Fengpeng Li, Kemou Li, Haiwei Wu, Jinyu Tian 0001, Jiantao Zhou 0001
IEEE Trans. Inf. Forensics Secur.4
2024 DifAttack: Query-Efficient Black-Box Adversarial Attack via Disentangled Feature Space
abstract
This work investigates efficient score-based black-box adversarial attacks with high Attack Success Rate (ASR) and good generalizability. We design a novel attack method based on a Disentangled Feature space, called DifAttack, which differs significantly from the existing ones operating over the entire feature space. Specifically, DifAttack firstly disentangles an image's latent feature into an adversarial feature and a visual feature, where the former dominates the adversarial capability of an image, while the latter largely determines its visual appearance. We train an autoencoder for the disentanglement by using pairs of clean images and their Adversarial Examples (AEs) generated from available surrogate models via white-box attack methods. Eventually, DifAttack iteratively optimizes the adversarial feature according to the query feedback from the victim model until a successful AE is generated, while keeping the visual feature unaltered. In addition, due to the avoidance of using surrogate models' gradient information when optimizing AEs for black-box models, our proposed DifAttack inherently possesses better attack capability in the open-set scenario, where the training dataset of the victim model is unknown. Extensive experimental results demonstrate that our method achieves significant improvements in ASR and query efficiency simultaneously, especially in the targeted attack and open-set scenarios. The code is available The code is available at https://github.com/csjunjun/DifAttack.git.
Jun Liu 0071, Jiantao Zhou 0001, Jiandian Zeng, Jinyu Tian 0001
AAAI4
2024 Regroup Median Loss for Combating Label Noise
abstract
The deep model training procedure requires large-scale datasets of annotated data. Due to the difficulty of annotating a large number of samples, label noise caused by incorrect annotations is inevitable, resulting in low model performance and poor model generalization. To combat label noise, current methods usually select clean samples based on the small-loss criterion and use these samples for training. Due to some noisy samples similar to clean ones, these small-loss criterion-based methods are still affected by label noise. To address this issue, in this work, we propose Regroup Median Loss (RML) to reduce the probability of selecting noisy samples and correct losses of noisy samples. RML randomly selects samples with the same label as the training samples based on a new loss processing method. Then, we combine the stable mean loss and the robust median loss through a proposed regrouping strategy to obtain robust loss estimation for noisy samples. To further improve the model performance against label noise, we propose a new sample selection strategy and build a semi-supervised method based on RML. Compared to state-of-the-art methods, for both the traditionally trained and semi-supervised models, RML achieves a significant improvement on synthetic and complex real-world datasets. The source is at https://github.com/Feng-peng-Li/Regroup-Loss-Median-to-Combat-Label-Noise.
Fengpeng Li, Kemou Li, Jinyu Tian 0001, Jiantao Zhou 0001
AAAI3
2024 DAT: Improving Adversarial Robustness via Generative Amplitude Mix-up in Frequency Domain
abstract
To protect deep neural networks (DNNs) from adversarial attacks, adversarial training (AT) is developed by incorporating adversarial examples (AEs) into model training. Recent studies show that adversarial attacks disproportionately impact the patterns within the phase of the sample's frequency spectrum---typically containing crucial semantic information---more than those in the amplitude, resulting in the model's erroneous categorization of AEs. We find that, by mixing the amplitude of training samples' frequency spectrum with those of distractor images for AT, the model can be guided to focus on phase patterns unaffected by adversarial perturbations. As a result, the model's robustness can be improved. Unfortunately, it is still challenging to select appropriate distractor images, which should mix the amplitude without affecting the phase patterns. To this end, in this paper, we propose an optimized **Adversarial Amplitude Generator (AAG)** to achieve a better tradeoff between improving the model's robustness and retaining phase patterns. Based on this generator, together with an efficient AE production procedure, we design a new **Dual Adversarial Training (DAT)** strategy. Experiments on various datasets show that our proposed DAT leads to significantly improved robustness against diverse adversarial attacks. The source code is available at https://github.com/Feng-peng-Li/DAT.
Fengpeng Li, Kemou Li, Haiwei Wu, Jinyu Tian 0001, Jiantao Zhou 0001
NeurIPS4
2024 Transformer-Based Image Inpainting Detection via Label Decoupling and Constrained Adversarial Training
abstract
Image inpainting based on generative adversarial networks (GANs) has achieved great success in producing visually plausible images and plays an important role in many real tasks. However, the techniques of image inpainting might also be maliciously used, e.g., altering or removing interesting objects to report fake news. Despite the promising performance of recently developed inpainting detection algorithms, they are built on convolutional neural networks (CNNs) with limited receptive fields. Consequently, they fail to fully capture the disparity between the inpainted regions and untouched regions and thus are ineffective in obtaining fine-grained detection results. In this work, we develop a new image inpainting detection approach. First, we propose a locally enhanced transformer architecture tailored for image inpainting detection. Unlike previous CNN-based methods, our approach leverages both the short-range and long-range dependencies of pixels, enabling the learning of diverse statistical behaviors of inpainted and untouched regions. Second, to mitigate the distraction caused by near-edge pixels with a mixed nature during training, we propose decoupling the label into a body map and a soft-edge map, and then a cross-modality attention module is designed to propagate their information interactively. It demonstrates that our decoupling strategy outperforms the conventional edge supervision in enhancing detection accuracy. Finally, we devise a constrained adversarial training methodology in consideration of the confrontational generation procedure of deep image inpainting methods. It shows that our constrained adversarial training further enhances the detection performance by adaptively introducing interference noise in the inpainted regions. Extensive experiments validate the superiority of our scheme compared to existing CNN-based methods, showcasing its desirable detection generalizability for both deep inpainting and traditional inpainting algorithms.
Yuanman Li, Liangpei Hu, Li Dong 0006, Haiwei Wu, Jinyu Tian 0001, Jiantao Zhou 0001, Xia Li 0006
IEEE Trans. Circuits Syst. Video Technol.5
2024 Robust Camera Model Identification Over Online Social Network Shared Images via Multi-Scenario Learning
abstract
Camera model identification (CMI) can be widely used in image forensics such as authenticity determination, copyright protection, forgery detection, etc. Meanwhile, with the vigorous development of the Internet, online social networks (OSNs) have become the dominant channels for image sharing and transmission. However, the inevitable lossy operations on OSNs, such as compression and post-processing, impose great challenges to the existing CMI schemes, as they severely destroy the camera traces left in the images under investigation. In this work, we propose a novel CMI method that is robust against the lossy operations of various OSN platforms. Specifically, it is observed that a camera trace extractor can be easily trained on a single degradation scenario (e.g., one specific OSN platform); while much more difficult on mixed degradation scenarios (e.g., multiple OSN platforms). Inspired by this observation, we design a new multi-scenario learning (MSL) strategy, enabling us to extract robust camera traces across different OSNs. Furthermore, noticing that image smooth regions incur less distortions by OSN and less interference by image signal itself, we suggest a SmooThness-Aware Trace Extractor (STATE) that can adaptively extract camera traces according to the smoothness of the input image. The superiority of our method is verified by comparative experiments with four state-of-the-art methods, especially under various OSN transmission scenarios. Particularly, for the open-set camera model verification task, we greatly surpass the second-place by 15.30% in AUC on theFODBdataset; while for the close-set camera model classification task, we are significantly ahead of the second-place by 34.51% in F1 on theSIHDRdataset. The code of our proposed method is available athttps://github.com/HighwayWu/CameraTraceOSN.
Haiwei Wu, Jiantao Zhou 0001, Jinyu Tian 0001, Weiwei Sun 0009
IEEE Trans. Inf. Forensics Secur.4
2024 Recoverable Privacy-Preserving Image Classification through Noise-like Adversarial Examples
abstract
With the increasing prevalence of cloud computing platforms, ensuring data privacy during the cloud-based image-related services such as classification has become crucial. In this study, we propose a novel privacy-preserving image classification scheme that enables the direct application of classifiers trained in the plaintext domain to classify encrypted images without the need of retraining a dedicated classifier. Moreover, encrypted images can be decrypted back into their original form with high fidelity (recoverable) using a secret key. Specifically, our proposed scheme involves utilizing a feature extractor and an encoder to mask the plaintext image through a newly designed Noise-like Adversarial Example (NAE). Such an NAE not only introduces a noise-like visual appearance to the encrypted image but also compels the target classifier to predict the ciphertext as the same label as the original plaintext image. At the decoding phase, we adopt a Symmetric Residual Learning (SRL) framework for restoring the plaintext image with minimal degradation. Extensive experiments demonstrate that (1) the classification accuracy of the classifier trained in the plaintext domain remains the same in both the ciphertext and plaintext domains; (2) the encrypted images can be recovered into their original form with an average PSNR of up to 51+ dB for the SVHN dataset and 48+ dB for the VGGFace2 dataset; (3) our system exhibits satisfactory generalization capability on the encryption, decryption, and classification tasks across datasets that are different from the training one; and (4) a high-level of security is achieved against three potential threat models. The code is available at https://github.com/csjunjun/RIC.git .
Jun Liu 0071, Jiantao Zhou 0001, Jinyu Tian 0001, Weiwei Sun 0009
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Generating Robust Adversarial Examples against Online Social Networks (OSNs)
abstract
Online Social Networks (OSNs) have blossomed into prevailing transmission channels for images in the modern era. Adversarial examples (AEs) deliberately designed to mislead deep neural networks (DNNs) are found to be fragile against the inevitable lossy operations conducted by OSNs. As a result, the AEs would lose their attack capabilities after being transmitted over OSNs. In this work, we aim to design a new framework for generating robust AEs that can survive the OSN transmission; namely, the AEs before and after the OSN transmission both possess strong attack capabilities. To this end, we first propose a differentiable network termed SImulated OSN (SIO) to simulate the various operations conducted by an OSN. Specifically, the SIO network consists of two modules: (1) a differentiable JPEG layer for approximating the ubiquitous JPEG compression and (2) an encoder-decoder subnetwork for mimicking the remaining operations. Based upon the SIO network, we then formulate an optimization framework to generate robust AEs by enforcing model outputs with and without passing through the SIO to be both misled. Extensive experiments conducted over Facebook, WeChat and QQ demonstrate that our attack methods produce more robust AEs than existing approaches, especially under small distortion constraints; the performance gain in terms of Attack Success Rate (ASR) could be more than 60%. Furthermore, we build a public dataset containing more than 10,000 pairs of AEs processed by Facebook, WeChat or QQ, facilitating future research in the robust AEs generation. The dataset and code are available at https://github.com/csjunjun/RobustOSNAttack.git .
Jun Liu 0071, Jiantao Zhou 0001, Haiwei Wu, Weiwei Sun 0009, Jinyu Tian 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Effective Ambiguity Attack Against Passport-based DNN Intellectual Property Protection Schemes through Fully Connected Layer Substitution
abstract
Since training a deep neural network (DNN) is costly, the well-trained deep models can be regarded as valuable intellectual property (IP) assets. The IP protection associated with deep models has been receiving increasing attentions in recent years. Passport-based method, which replaces normalization layers with passport layers, has been one of the few protection solutions that are claimed to be secure against advanced attacks. In this work, we tackle the issue of evaluating the security of passport-based IP protection methods. We propose a novel and effective ambiguity attack against passport-based method, capable of successfully forging multiple valid passports with a small training dataset. This is accomplished by inserting a specially designed accessory block ahead of the passport parameters. Using less than 10% of training data, with the forged passport, the model exhibits almost indistinguishable performance difference (less than 2%) compared with that of the authorized passport. In addition, it is shown that our attack strategy can be readily generalized to attack other IP protection methods based on watermark embedding. Directions for potential remedy solutions are also given.
Jinyu Tian 0001, Xiangyu Chen 0006, Jiantao Zhou 0001
CVPR2
2023 Gradient Sign Inversion: Making an Adversarial Attack a Good Defense
abstract
Deep neural networks have been proven vulnerable to deliberately crafted adversarial example, which cause serious safety and security concerns. Many defense approaches were proposed to resist such threats. However, existing defenses such as pre-compression or adversarial training would degrade the model performance on clean images or incur heavy computational costs. In this work, we propose a plug-and-play defensive module Gradient Sign Inversion (GSI) to defend gradient-based attack. Essentially, GSI attempts to inverse the direction of the backpropagated gradient for the victim model, disturbing the adversarial example generation of the attacking while retaining the performance of the vanilla network on genuine inputs. Specifically, an additive model based on periodic trigonometric function is established by investigating the necessary conditions that a suitable defensive module should have. By enforcing constraints on the defensive module, the parameters of GSI are determined, accompanied by a theoretical justification. Interestingly, we observe that the proposed GSI not only prevents the gradient-based adversarial attack, but can even improve the confidence of the ground-truth label when initiating an attack, making the attack betray as a defense. Source code is publicly available at https://github.com/JidaDiao/GSI.
Xiaojian Ji, Li Dong 0006, Rangding Wang, Diqun Yan, Yang Yin, Jinyu Tian 0001
IJCNN6
2023 Universal Defensive Underpainting Patch: Making Your Text Invisible to Optical Character Recognition
abstract
Optical Character Recognition (OCR) enables automatic text extraction from scanned or digitized text images, but it also makes it easy to pirate valuable or sensitive text from these images. Previous methods to prevent OCR piracy by distorting characters in text images are impractical in real-world scenarios, as pirates can capture arbitrary portions of the text images, rendering the defenses ineffective. In this work, we propose a novel and effective defense mechanism termed the Universal Defensive Underpainting Patch (UDUP) that modifies the underpainting of text images instead of the characters. UDUP is created through an iterative optimization process to craft a small, fixed-size defensive patch that can generate non-overlapping underpainting for text images of any size. Experimental results show that UDUP effectively defends against unauthorized OCR under the setting of any screenshot range or complex image background. It is agnostic to the content, size, colors, and languages of characters, and is robust to typical image operations such as scaling and compressing. In addition, the transferability of UDUP is demonstrated by evading several off-the-shelf OCRs. The code is available at https://github.com/QRICKDD/UDUP.
Jiacheng Deng 0001, Li Dong 0006, Diqun Yan, Rangding Wang, Dengpan Ye, Lingchen Zhao, Jinyu Tian 0001
ACM Multimedia8
2023 Parallel multiple watermarking using adaptive Inter-Block correlation
Xingrun Wang, Xiaochen Yuan, Mianjie Li, Jinyu Tian 0001, Hongfei Guo, Jianqing Li 0001
Expert Syst. Appl.5
2023 Microcontroller Unit Chip Temperature Fingerprint Informed Machine Learning for IIoT Intrusion Detection
abstract
Physics-informed learning for industrial Internet is essential especially to safety issues. Consequently, various methods have been developed to conduct Industrial Internet of Things (IIoT) intrusion detection. However, the conventional methods usually require the help of auxiliary equipment (e.g., spectrum analyzers, log-periodic antennas), which proves to be unsuitable for general IIoT systems due to their poor versatility. Facing the dilemma mentioned above, this article proposes a microcontroller unit (MCU) chip temperature fingerprint informed machine learning method, called MTID, for IIoT intrusion detection. Specifically, first, the node's MCU temperature sequence is recorded and the relationship between the temperature sequence and the computational complexity of the node is analyzed. Then, we calculate the temperature residuals and construct a temperature residuals dataset. Finally, to identify the security status of the nodes, a self-encoder-based intrusion detection model is constructed. Furthermore, to ensure the model's applicability under the diversified deployment environment of IIoT systems, an online incremental training method is developed and applied. In the end, we use the Raspberry Pi 4B for experimental analysis when testing the performance of MTID. The results show that the accuracy of MTID for intrusion detection reaches 89%, which also demonstrates the feasibility of the intrusion detection method based on MCU temperature.
Tingting Wang 0006, Kai Fang 0001, Wei Wei 0006, Jinyu Tian 0001, Yuanyuan Pan, Jianqing Li 0001
IEEE Trans. Ind. Informatics4
2023 Spatial-Temporal Attention Graph Convolution Network on Edge Cloud for Traffic Flow Prediction
abstract
Accurate short-term traffic flow prediction plays an important role in providing road condition information in the immediate future. With the information, intelligent vehicles can plan and adjust the route to prevent congestion. As a result, many models for short-term traffic flow forecasting have been proposed to date. However, most of them focus on the prediction of the entire traffic network, which could lead to several problems: (1) the entire traffic network could have a large scale and a complex structure, for which the model training is likely to be time-consuming as well as inefficient; (2) processing a large amount of training data on the central cloud could cause much calculation pressure on the server and increase the risk of privacy leakage. In this paper, we propose a Spatial-Temporal Attention Graph Convolution Network on Edge Cloud model (STAGCN-EC). We first divide the entire traffic network into several parts to reduce its scale and complexity. Then, we allocate each part of the network to a certain Roadside Unit (RSU) for training, thus there is no need to process all data on the central server. Besides, we utilize spatial-temporal attention and features extracting module that fits the low computational power devices like RSUs, to capture spatial-temporal dependence and predict traffic flow. At last, we use two highway datasets from District 7 and District 4 in California to validate our model. Through the experiments, we find out that our model performs well both in predicted precision and efficiency compared with the five baseline methods.
Qifeng Lai, Jinyu Tian 0001, Wei Wang 0077, Xiping Hu
IEEE Trans. Intell. Transp. Syst.2
2023 Hierarchical Services of Convolutional Neural Networks via Probabilistic Selective Encryption
abstract
Model protection is vital when deploying Convolutional Neural Networks (CNNs) for commercial services, due to the massive costs of training them. In this work, we propose a selective encryption (SE) algorithm to protect CNN models from unauthorized access, with a unique feature of providing hierarchical services to users. Our algorithm firstly selects important model parameters via the proposed Probabilistic Selection Strategy (PSS). It then encrypts the most important parameters with the designed encryption method called Distribution Preserving Random Mask (DPRM), so as to maximize the performance degradation by encrypting only a very small portion of model parameters. We also design a set of access permissions, using which different amount of most important model parameters can be decrypted. Hence, different levels of model performance can be naturally provided for users. Experimental results demonstrate that the proposed scheme could effectively protect the classification model VGG19 by merely encrypting 8% parameters of convolutional layers. We also implement the proposed model protection scheme in the denoising model DnCNN, showcasing the hierarchical denoising services.
Jinyu Tian 0001, Jiantao Zhou 0001, Jia Duan
IEEE Trans. Serv. Comput.1
2022 Robust Image Forgery Detection over Online Social Network Shared Images
abstract
The increasing abuse of image editing softwares, such as Photoshop and Meitu, causes the authenticity of digital images questionable. Meanwhile, the widespread availability of online social networks (OSNs) makes them the dominant channels for transmitting forged images to report fake news, propagate rumors, etc. Unfortunately, various lossy operations adopted by OSNs, e.g., compression and resizing, impose great challenges for implementing the robust image forgery detection. To fight against the OSN-shared forgeries, in this work, a novel robust training scheme is proposed. We first conduct a thorough analysis of the noise introduced by OSNs, and decouple it into two parts, i.e., predictable noise and unseen noise, which are modelled separately. The former simulates the noise introduced by the disclosed (known) operations of OSNs, while the latter is designed to not only complete the previous one, but also take into account the defects of the detector itself. We then incorporate the modelled noise into a robust training framework, significantly improving the robustness of the image forgery detector. Extensive experimental results are presented to validate the superiority of the proposed scheme compared with several state-of-the-art competitors. Finally, to promote the future development of the image forgery detection, we build a public forgeries dataset based on four existing datasets and three most popular OSNs. The designed detector recently won the top ranking in a certificate forgery detection competition11https://tianchi.aliyun.com/competition/entrance/531812/introduction. The source code and dataset are available at https://github.com/HighwayWu/lmageForensicsOSN.
Haiwei Wu, Jiantao Zhou 0001, Jinyu Tian 0001, Jun Liu 0071
CVPR3
2022 Self-Supervised Adversarial Example Detection by Disentangled Representation
abstract
Deep learning models are known to be vulnerable to adversarial examples that are elaborately designed for malicious purposes and are imperceptible to the human perceptual system. Autoencoder, when trained solely over benign examples, has been widely used for (self-supervised) adversarial detection based on the assumption that adversarial examples yield larger reconstruction errors. However, because lacking adversarial examples in its training and the too strong generalization ability of autoencoder, this assumption does not always hold true in practice. To alleviate this problem, we explore how to detect adversarial examples with disentangled label/semantic features under the autoencoder structure. Specifically, we propose Disentangled Representation-based Reconstruction (DRR). In DRR, we train an autoencoder over both correctly paired label/semantic features and incorrectly paired label/semantic features to reconstruct benign and counterexamples. This mimics the behavior of adversarial examples and can reduce the unnecessary generalization ability of autoencoder. We compare our method with the state-of-the-art self-supervised detection methods under different adversarial attacks and different victim models, and it exhibits better performance in various metrics (area under the ROC curve, true positive rate, and true negative rate) for most attack settings. Though DRR is initially designed for visual tasks only, we demonstrate that it can be easily extended for natural language tasks as well. Notably, different from other autoencoder-based detectors, our method can provide resistance to the adaptive adversary.
Zhaoxi Zhang 0001, Leo Yu Zhang, Xufei Zheng, Jinyu Tian 0001, Jiantao Zhou 0001
TrustCom4
2022 Robust Matrix Factorization via Minimum Weighted Error Entropy Criterion
abstract
Learning the intrinsic low-dimensional subspace from high-dimensional data is a key step for many social systems of artificial intelligence. In practical scenarios, the observed data are usually corrupted by many types of noise, which brings a great challenge for social systems to analyze data. As a commonly utilized subspace learning technique, robust low-rank matrix factorization (LRMF) focuses on recovering the underlying subspaces in a noisy environment. However, most of the existing approaches simply assume that the noise contaminating the data is independent identically distributed (i.i.d.), such as Gaussian and Laplacian noises. This assumption, though greatly simplifies the underlying learning problem, may not hold for more complex non-i.i.d. noise widely existed in social systems. In this work, we suggest a robust LRMF approach to deal with various types of noise in a unified manner. Different from traditional algorithms, noise in our framework is modeled using an independent and piecewise identically distributed (i.p.i.d.) source, which employs a collection of distributions, instead of a single one to characterize the statistical behavior of the underlying noise. Assisted by the generic noise model, we then design a robust LRMF algorithm under the information-theoretic learning (ITL) framework through a new minimization criterion. By adopting the half-quadratic optimization paradigm, we further deliver an optimization strategy for our proposed method. Experimental results on both synthetic and real data are provided to demonstrate the superiority of our proposed scheme.
Yuanman Li, Jiantao Zhou 0001, Junyang Chen 0001, Jinyu Tian 0001, Li Dong 0006, Xia Li 0006
IEEE Trans. Comput. Soc. Syst.4
2022 Robust Image Forgery Detection Against Transmission Over Online Social Networks
abstract
The increasing abuse of image editing software causes the authenticity of digital images questionable. Meanwhile, the widespread availability of online social networks (OSNs) makes them the dominant channels for transmitting forged images to report fake news, propagate rumors, etc. Unfortunately, various lossy operations, e.g., compression and resizing, adopted by OSNs impose great challenges for implementing the robust image forgery detection. To fight against the OSN-shared forgeries, in this work, a novel robust training scheme is proposed. Firstly, we design a baseline detector, which won the top ranking in a recent certificate forgery detection competition. Then we conduct a thorough analysis of the noise introduced by OSNs, and decouple it into two parts, i.e.,predictable noiseandunseen noise, which are modelled separately. The former simulates the noise introduced by the disclosed (known) operations of OSNs, while the latter is designed to not only complete the previous one, but also take into account the defects of the detector itself. We further incorporate the modelled noise into a robust training framework, significantly improving the robustness of the image forgery detector. Extensive experimental results are presented to validate the superiority of the proposed scheme compared with several state-of-the-art competitors, especially in the scenarios of detecting OSN-transmitted forgeries. Finally, to promote the future development of the image forgery detection, we build a public forgeries dataset based on four existing datasets through the uploading and downloading of four most popular OSNs. The data and code of this work are available athttps://github.com/HighwayWu/ImageForensicsOSN.
Haiwei Wu, Jiantao Zhou 0001, Jinyu Tian 0001, Jun Liu 0071, Yu Qiao 0001
IEEE Trans. Inf. Forensics Secur.3
2022 Weighted Error Entropy-Based Information Theoretic Learning for Robust Subspace Representation
abstract
In most of the existing representation learning frameworks, the noise contaminating the data points is often assumed to be independent and identically distributed (i.i.d.), where the Gaussian distribution is often imposed. This assumption, though greatly simplifies the resulting representation problems, may not hold in many practical scenarios. For example, the noise in face representation is usually attributable to local variation, random occlusion, and unconstrained illumination, which is essentially structural, and hence, does not satisfy the i.i.d. property or the Gaussianity. In this article, we devise a generic noise model, referred to as independent and piecewise identically distributed (i.p.i.d.) model for robust presentation learning, where the statistical behavior of the underlying noise is characterized using a union of distributions. We demonstrate that our proposed i.p.i.d. model can better describe the complex noise encountered in practical scenarios and accommodate the traditional i.i.d. one as a special case. Assisted by the proposed noise model, we then develop a new information-theoretic learning framework for robust subspace representation through a novel minimum weighted error entropy criterion. Thanks to the superior modeling capability of the i.p.i.d. model, our proposed learning method achieves superior robustness against various types of noise. When applying our scheme to the subspace clustering and image recognition problems, we observe significant performance gains over the existing approaches.
Yuanman Li, Jiantao Zhou 0001, Jinyu Tian 0001, Xianwei Zheng, Yuan Yan Tang
IEEE Trans. Neural Networks Learn. Syst.3
2021 Detecting Adversarial Examples from Sensitivity Inconsistency of Spatial-Transform Domain
abstract
Deep neural networks (DNNs) have been shown to be vulnerable against adversarial examples (AEs), which are maliciously designed to cause dramatic model output errors. In this work, we reveal that normal examples (NEs) are insensitive to the fluctuations occurring at the highly-curved region of the decision boundary, while AEs typically designed over one single domain (mostly spatial domain) exhibit exorbitant sensitivity on such fluctuations. This phenomenon motivates us to design another classifier (called dual classifier) with transformed decision boundary, which can be collaboratively used with the original classifier (called primal classifier) to detect AEs, by virtue of the sensitivity inconsistency. When comparing with the state-of-the-art algorithms based on Local Intrinsic Dimensionality (LID), Mahalanobis Distance (MD), and Feature Squeezing (FS), our proposed Sensitivity Inconsistency Detector (SID) achieves improved AE detection performance and superior generalization capabilities, especially in the challenging cases where the adversarial perturbation levels are small. Intensive experimental results on ResNet and VGG validate the superiority of the proposed SID.
Jinyu Tian 0001, Jiantao Zhou 0001, Yuanman Li, Jia Duan
AAAI1
2021 Probabilistic Selective Encryption of Convolutional Neural Networks for Hierarchical Services
abstract
Model protection is vital when deploying Convolutional Neural Networks (CNNs) for commercial services, due to the massive costs of training them. In this work, we propose a selective encryption (SE) algorithm to protect CNN models from unauthorized access, with a unique feature of pro-viding hierarchical services to users. Our algorithm firstly selects important model parameters via the proposed Probabilistic Selection Strategy (PSS). It then encrypts the most important parameters with the designed encryption method called Distribution Preserving Random Mask (DPRM), so as to maximize the performance degradation by encrypting only a very small portion of model parameters. We also design a set of access permissions, using which different amount of most important model parameters can be decrypted. Hence, different levels of model performance can be naturally provided for users. Experimental results demonstrate that the proposed scheme could effectively protect the classification model VGG19 by merely encrypting 8% parameters of convolutional layers. We also implement the proposed model protection scheme in the denoising model DnCNN, showcasing the hierarchical denoising services.
Jinyu Tian 0001, Jiantao Zhou 0001, Jia Duan
CVPR1
2021 Optimal Pre-Filtering for Improving Facebook Shared Images
abstract
Online Social Networks (OSNs) have attracted a huge number of users, who store and share various images on a daily basis. As a well-known fact, most OSN platforms apply a series of lossy operations on the uploaded images, which could severely degrade the quality of the shared images, negatively affecting the user experiences. In this work, we consider the problem of significantly improving OSN-shared images through applying an optimal pre-filtering prior to image sharing, without any cooperation from the OSN platform itself. Facebook, as one of the most popular and representative OSNs, is chosen as the platform to present our designed pre-filtering strategy. We first treat Facebook as a black box, and thoroughly recover its mechanism of processing color images. Based on the precise knowledge on the image processing pipeline on Facebook, we design the pre-filter under an optimization framework, minimizing the end-to-end distortion between the shared image and the original one. Compared with the directly shared images, our proposed pre-filtering-then-sharing strategy brings significant improvements in terms of both quantitative and qualitative metrics. Extensive experimental results are provided to show the superiority of our proposed method. Finally, we discuss the strategy on how to extend our proposed technique to other OSN platforms.
Weiwei Sun 0009, Jiantao Zhou 0001, Li Dong 0006, Jinyu Tian 0001, Jun Liu 0071
IEEE Trans. Image Process.4
2019 Robust Subspace Clustering With Independent and Piecewise Identically Distributed Noise Modeling
abstract
Most of the existing subspace clustering (SC) frameworks assume that the noise contaminating the data is generated by an independent and identically distributed (i.i.d.) source, where the Gaussianity is often imposed. Though these assumptions greatly simplify the underlying problems, they do not hold in many real-world applications. For instance, in face clustering, the noise is usually caused by random occlusions, local variations and unconstrained illuminations, which is essentially structural and hence satisfies neither the i.i.d. property nor the Gaussianity. In this work, we propose an independent and piecewise identically distributed (i.p.i.d.) noise model, where the i.i.d. property only holds locally. We demonstrate that the i.p.i.d. model better characterizes the noise encountered in practical scenarios, and accommodates the traditional i.i.d. model as a special case. Assisted by this generalized noise model, we design an information theoretic learning (ITL) framework for robust SC through a novel minimum weighted error entropy (MWEE) criterion. Extensive experimental results show that our proposed SC scheme significantly outperforms the state-of-the-art competing algorithms.
Yuanman Li, Jiantao Zhou 0001, Xianwei Zheng, Jinyu Tian 0001, Yuan Yan Tang
CVPR4
2019 Spectral-Spatial Graph Convolutional Networks for Semisupervised Hyperspectral Image Classification
abstract
Collecting labeled samples is quite costly and time-consuming for hyperspectral image (HSI) classification task. Semisupervised learning framework, which combines the intrinsic information of labeled and unlabeled samples, can alleviate the deficient labeled samples and increase the accuracy of HSI classification. In this letter, we propose a novel semisupervised learning framework that is based on spectral-spatial graph convolutional networks (S2GCNs). It explicitly utilizes the adjacency nodes in graph to approximate the convolution. In the process of approximate convolution on graph, the proposed method makes full use of the spatial information of the current pixel. The experimental results on three real-life HSI data sets, i.e., Botswana Hyperion, Kennedy Space Center, and Indian Pines, show that the proposed S2GCN can significantly improve the classification accuracy. For instance, the overall accuracy on Indian data is increased from 66.8% (GCN) to 91.6%.
Anyong Qin, Zhaowei Shang, Jinyu Tian 0001, Yulong Wang 0002, Taiping Zhang, Yuan Yan Tang
IEEE Geosci. Remote. Sens. Lett.3
2017 Maximum correntropy criterion for convex anc semi-nonnegative matrix factorization
abstract
Matrix factorization is a popular low dimensional representation approach that plays an important role in many pattern recognition and computer vision domains. Among them, convex and semi-nonnegative matrix factorizations have attracted considerable interest, owing to its clustering interpretation. On the other hand, the generalized correlation function (correntropy) as the error measure does not depend on the assumption of Gaussianity, which the mean square error (MSE) heavily depends on. In this paper, we propose two novel algorithms, called Maximum Correntropy Criterion based Convex and Semi-Nonnegative Matrix Factorization (MCC-ConvexNMF, MCC-SemiNMF). Compared with the mean square error based convex and semi-nonnegative matrix factorization, the proposed methods can extract more information from the data and produce more accurate solutions. Experimental results on both synthetic dataset and the popular face database illustrate the effectiveness of our methods.
Anyong Qin, Zhaowei Shang, Jinyu Tian 0001, Ailin Li, Yulong Wang 0002, Yuan Yan Tang
SMC3
2017 Learning the Distribution Preserving Semantic Subspace for Clustering
abstract
This paper proposes a new clustering method for images called distribution preserving indexing (DPI). It aims to find a lower dimensional semantic space approximating the original image space in the sense of preserving the distribution of the data. In the theory, the intrinsic structure of the data clusters can be described by the distribution of the data effectively. Therefore, the cluster structure of the data in a lower dimensional semantic space derived by the DPI becomes clear. Unlike these distance-based clustering methods, which reveal the intrinsic Euclidean structure of data, our method attempts to discover the intrinsic cluster structure of the data space that actually is the union of some sub-manifolds. Moreover, we propose a revised kernel density estimator for the case of high-dimensional data, which is a crucial step in DPI. In addition, we provide a theoretical analysis of the bound of our method. Finally, the extensive experiments compared with other algorithms, on COIL20, CBCL, and MNIST demonstrate the effectiveness of our proposed approach.
Jinyu Tian 0001, Taiping Zhang, Anyong Qin, Zhaowei Shang, Yuan Yan Tang
IEEE Trans. Image Process.1