Junxian Duan

dblp:187/6138 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 TVChain: Leveraging Textual-Visual Prompt Chains for Jailbreaking Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) enhance the capabilities of Large Language Models by integrating visual inputs, thereby enabling advanced multimodal reasoning across diverse applications. However, these enhanced reasoning capabilities introduce new security risks, particularly to jailbreaking attacks that bypass built-in safety mechanisms to elicit harmful or unauthorized outputs. While recent efforts have explored adversarial and typographic prompts, most existing attacks suffer from three key limitations: reliance on auxiliary models, limited effectiveness in black-box scenarios, and inadequate exploitation of the LVLMs' intrinsic reasoning abilities. In this work, we propose TVChain, a novel black-box jailbreaking framework that explicitly intervenes in both the visual and textual reasoning processes of LVLMs. TVChain decomposes malicious prompts into a sequence of semantically meaningful sub-images that represent relevant objects and behaviors, thereby circumventing direct exposure of illicit content. In parallel, a carefully designed chain-of-thought (CoT) textual prompt is employed to steer the model's reasoning toward reconstructing the intended activity in a covert yet effective manner. We demonstrate that this compositional prompting strategy reduces the likelihood of triggering safety mechanisms while preserving attack efficacy. Extensive evaluations on eleven LVLMs (seven open-source and four commercial) across two benchmark datasets and three state-of-the-art defenses validate the effectiveness and robustness of TVChain.
Hao Yu 0017, Ke Liang 0006, Junxian Duan, Jun Wang 0118, Siwei Wang 0001, Chuan Ma 0001, Xinwang Liu 0002
AAAI3
2025 Dual-PST: Dual-Branch SpatioTemporal-Planar Network for Video Forgery Detection
abstract
With the advancement of generative AI, distinguishing real and AI-generated faces in videos has become increasingly challenging. However, traditional methods struggle to capture local details and temporal dynamics simultaneously, making it difficult to achieve high detection accuracy while maintaining low computational overhead. To address this problem, we propose a Dual-branch SpatioTemporal-Planar Network (Dual-PST) based on the selective state-space model. It is capable of extracting image features and temporal relations simultaneously, while maintaining linear computational consumption. Specifically, we design a Multi-Selective State-Space module (MS3) that can extract global features from image typography consisting of consecutive video frames by scanning them in multiple sequences. To further enhance temporal modeling capabilities, we propose a Sequential Tri-frame Local module, which captures inter-frame temporal relationships and local features by temporally splicing single-frame features. These features are first extracted using MS3 and then further enhanced through inter-frame masking operations. Experimental results show that Dual-PST significantly improves detection accuracy while maintaining low computational complexity and strong model robustness.
Junxian Duan, Jie Cao 0002, Aihua Zheng
ICASSP3
2025 Growing to Detect: A Dynamic Prototype Tree with Structured Replay for Incremental Deepfake Detection
abstract
The rapid advancement of deepfake technology poses significant threats to social trust. Recent research has improved detectors by adapting to emerging deepfakes using a limited number of samples through incremental learning. However, these approaches often overlook the scarcity of novel samples, resulting in insufficient learning of forgery patterns. To overcome this challenge, we propose a Replay-based Dynamic Prototype Network that integrates two key modules: the Dynamic Prototype Tree (DPT) module and the Similarity Subtree Replay (SSR) strategy. The DPT module dynamically introduces prototypes through a hierarchical tree structure to effectively adapt to new deepfakes. It expands prototypes based on similarity, thereby retaining the knowledge learned from previous prototypes while learning new forgery patterns. The SSR strategy mitigates catastrophic forgetting by stabilizing learned features through the replay of relevant subtrees. Experimental results demonstrate that our approach outperforms existing methods across five datasets, particularly on high-quality face swap samples generated by diffusion-based methods, achieving an AUC of 85.86% on the cross-dataset task from FaceForensics++ to DiffSwap.
Junxian Duan, Jie Cao 0002, Aihua Zheng, Ran He 0001
IJCB2
2025 Towards Robust Defense Against Customization via Protective Perturbation Resistant to Diffusion-based Purification
abstract
Diffusion models like Stable Diffusion have become prominent in visual synthesis tasks due to their powerful customization capabilities, which also introduce significant security risks, including deepfakes and copyright infringement. In response, a class of methods known as protective perturbation emerged, which mitigates image misuse by injecting imperceptible adversarial noise. However, purification can remove protective perturbations, thereby exposing images again to the risk of malicious forgery. In this work, we formalize the anti-purification task, highlighting challenges that hinder existing approaches, and propose a simple diagnostic protective perturbation named AntiPure. AntiPure exposes vulnerabilities of purification within the "purification-customization" workflow, owing to two guidance mechanisms: 1) Patch-wise Frequency Guidance, which reduces the model's influence over high-frequency components in the purified image, and 2) Erroneous Timestep Guidance, which disrupts the model's denoising strategy across different timesteps. With additional guidance, AntiPure embeds imperceptible perturbations that persist under representative purification settings, achieving effective post-customization distortion. Experiments show that, as a stress test for purification, AntiPure achieves minimal perceptual discrepancy and maximal distortion, outperforming other protective perturbation methods within the purification-customization workflow.
Wenkui Yang, Jie Cao 0002, Junxian Duan, Ran He 0001
ICCV3
2025 MTSD: Simple Yet Effective Self-Distillation for Generalizable Deepfake Detection
abstract
The rapid advancement of Deepfake technology necessitates detection systems with strong generalization capabilities. Existing methods often depend on architectural modifications or dataset-specific prior knowledge, which limits their scalability and practical deployment in real-world scenarios. We propose Multi-Teacher Self-Distillation (MTSD), a simple yet effective and generalizable strategy to enhance model generalization. MTSD comprises two key steps. First, diverse teacher generation leverages independently trained teacher models with varying dataset sampling sequences to capture complementary decision boundaries. Second, self-distillation feature fusion integrates these diverse features using a cross-attention mechanism, allowing the student model to approximate an ideal feature distribution for improved generalization. This strategy avoids architectural changes and dataset-specific adjustments, ensuring simplicity in implementation and deployment. Moreover, the multi-teacher generation and feature fusion steps are discarded after training, preserving computational efficiency during inference. Experimental results demonstrate that MTSD significantly improves model generalization, offering a practical and scalable solution for Deepfake detection.
Dexu Zhu, Jie Cao 0002, Jiangnan Shao, Junxian Duan, Ran He 0001
ICME5
2025 Text-Guided Noise Replacement Visual Prompt Learning for Vision-Language Models
Xiaokang Shao, Mengjin Liu, Zhaojun Liu, Junxian Duan, Aihua Zheng
PRCV (2)5
2025 Trustworthy forgery detection with causal inference
Junxian Duan, Fan Ji, Yi Li 0018, Ran He 0001
Sci. China Inf. Sci.1
2025 Test-time Forgery Detection with Spatial-Frequency Prompt Learning
Junxian Duan, Yuang Ai, Shenyuan Huang, Huaibo Huang, Jie Cao 0002, Ran He 0001
Int. J. Comput. Vis.1
2025 RealDTT: Towards A Comprehensive Real-World Dataset for Tampered Text Detection
Junxian Duan, Fan Ji, Zhiyong Wang 0001, Huaibo Huang
Int. J. Comput. Vis.1
2024 PortraitDAE: Line-Drawing Portraits Style Transfer from Photos via Diffusion Autoencoder with Meaningful Encoded Noise
abstract
The line-drawing portrait is a kind of highly abstract art that contains a sparse set of continuous graphical elements such as lines to capture a person's facial features. Due to their abstract artistic form, common style transfer methods fail to synthesize high-quality line-drawing portraits from photos. Previous works mostly concentrate on GANs, often requiring pre-calculated landmarks acquired by other models and using extra classifiers with complicated structures to capture local facial features. We propose a novel idea without these extra operations based on diffusion models, which is more flexible and stable than GAN-based methods. We utilize the diffusion-based decoder in the Diffusion Autoencoder to encode the input image to an encoded noise that contains much meaningful stochastic information by running the deterministic generative process backward. By fully utilizing the encoded noise, our method can effectively preserve the identity information and better capture facial details. We also improve the loss function to alleviate the interference of the background color. Several experiments show that our method can produce better samples with smoother lines that look more like the corresponding person, outperforming state-of-the-art methods both qualitatively and quantitatively. Our method can also be generalized to other styles such as sketch.
Yexiang Liu, Jin Liu 0040, Jie Cao 0002, Junxian Duan, Ran He 0001
FG4
2024 TT-DF: A Large-Scale Diffusion-Based Dataset and Benchmark for Human Body Forgery Detection
Wenkui Yang, Xiaoqiang Zhou, Junxian Duan, Jie Cao 0002
PRCV (11)4
2023 From Ledger to P2P Network: De-anonymization on Bitcoin Using Cross-Layer Analysis
Che Zheng, Meng Shen 0001, Junxian Duan, Liehuang Zhu
APPT3
2023 Where to Focus: Central Attention-Based Face Forgery Detection
Jinghui Sun, Yuhe Ding, Jie Cao 0002, Junxian Duan, Aihua Zheng
PRCV (5)4
2023 Iterative embedding distillation for open world vehicle recognition
Junxian Duan, Xiang Wu 0001, Yibo Hu 0001, Chaoyou Fu, Zi Wang 0013, Ran He 0001
Pattern Recognit.1
2021 Visual-Semantic Transformer for Face Forgery Detection
abstract
This paper proposes a novel Visual-Semantic Transformer (VST) to detect face forgery based on semantic aware feature relations. In face images, intrinsic feature relations exist between different semantic parsing regions. We find that face forgery algorithms always change such relations. Therefore, we start the approach by extracting Contextual Feature Sequence (CFS) using a transformer encoder to make the best abnormal feature relation patterns. Meanwhile, images are segmented as soft face regions by a face parsing module. Then we merge the CFS and the soft face regions as Visual Semantic Sequences (VSS) representing features of semantic regions. The VSS is fed into the transformer decoder, in which the relations in the semantic region level are modeled. Our method achieved 99.58% accuracy on FF++(Raw) and 96.16% accuracy on Celeb-DF. Extensive experiments demonstrate that our framework outperforms or is comparable with state-of-the-art detection methods, especially towards unseen forgery methods.
Gengyun Jia, Huaibo Huang, Junxian Duan, Ran He 0001
IJCB4
2020 Blockchain-Based Incentives for Secure and Collaborative Data Sharing in Multiple Clouds
abstract
The prosperity of cloud computing has driven an increasing number of enterprises and organizations to store their data on private or public cloud platforms. Due to the limitation of individual data owners in terms of data volume and diversity, data sharing over different cloud platforms would enable third parties to take advantage of big data analysis techniques to provide value-added services, such as providing healthcare services for customers by gathering medical data from multiple hospitals. However, it remains a challenging task to design effective incentives that encourage secure and collaborative data sharing in multiple clouds. In this paper, we propose a reliable collaboration model consisting of three types of participants, which include data owners, miners, and third parties, where the data is shared via blockchain and recorded by a smart contract. In general, these participants may acquire and store the sharing of data using their private or public clouds. We analyze the topological relationships between the participants and develop some Shapley value models from simple to complicate in the process of revenue distribution. We also discuss the incentive effect of sharing security data and rationality of the designed solution through analysis towards distribution rules.
Meng Shen 0001, Junxian Duan, Liehuang Zhu, Jie Zhang 0061, Xiaojiang Du, Mohsen Guizani
IEEE J. Sel. Areas Commun.2