EDBT 2026 Demo / reviewers in the wild / expert
Xun Lin
dblp:57/88
· DBLP profile ↗
35ranked-venue papers
4as first author
33since 2021 · last 2026
0000-0001-8387-4245ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 16 since 2021Security and privacy · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | StegaVAR: Privacy-Preserving Video Action Recognition via Steganographic Domain AnalysisabstractDespite the rapid progress of deep learning in video action recognition (VAR) in recent years, privacy leakage in videos remains a critical concern. Current state-of-the-art privacy-preserving methods often rely on anonymization. These methods suffer from (1) low concealment, where producing visually distorted videos that attract attackers’ attention during transmission, and (2) spatiotemporal disruption, where degrading essential spatiotemporal features for accurate VAR. To address these issues, we propose StegaVAR, a novel framework that embeds action videos into ordinary cover videos and directly performs VAR in the steganographic domain for the first time. Throughout both data transmission and action analysis, the spatiotemporal information of hidden secret video remains complete, while the natural appearance of cover videos ensures the concealment of transmission. Considering the difficulty of steganographic domain analysis, we propose Secret Spatio-Temporal Promotion (STeP) and Cross-Band Difference Attention (CroDA) for analysis within the steganographic domain. STeP uses the secret video to guide spatiotemporal feature extraction in the steganographic domain during training. CroDA suppresses cover interference by capturing cross-band semantic differences. Experiments demonstrate that StegaVAR achieves superior VAR and privacy-preserving performance on widely used datasets. Moreover, our framework is effective for multiple steganographic models. Lixin Chen, Chaomeng Chen, Zhijian Wu, Xun Lin |
AAAI | 5 |
| 2026 | PA-FAS: Towards Interpretable and Generalizable Multimodal Face Anti-Spoofing via Path-Augmented Reinforcement LearningabstractIn recent years, face anti-spoofing (FAS) has made notable progress in multimodal fusion, cross-domain generalization, and interpretability. With the development of large language models and reinforcement learning (RL), strategy-based training paradigms offer new opportunities for jointly modeling multimodality, generalization, and interpretability. However, compared to unimodal reasoning, multimodal reasoning introduces more complex logic, such as accurate feature representation and cross-modal verification, which significantly increases reasoning complexity and labeling difficulty. Due to the lack of high-quality annotations in existing multimodal FAS datasets, directly applying RL strategies is sub-optimal, hindering robust multimodal reasoning. In this paper, we find two key issues of supervised fine-tuning combined with reinforcement learning (SFT+RL) paradigms in multimodal FAS reasoning: 1) limited multimodal reasoning paths not only hinder the full utilization of multimodal information but also constrain the model’s exploration space after SFT, thereby affecting the effectiveness of subsequent RL; and 2) the mismatch between single-task supervision and the diversity of multimodal reasoning paths leads to reasoning confusion, where models may exploit shortcuts by directly mapping input images to answers, bypassing the intended reasoning process. These issues further increase the complexity of multimodal reasoning and hinder the effective application of RL strategies. To address these challenges, we propose the PA-FAS framework with a reasoning path enhancement strategy for high-quality extended reasoning sequences construction based on limited annotated data to enrich the reasoning paths and alleviate exploration constraints. Additionally, we introduce an answer shuffling mechanism during SFT for comprehensive multimodal analysis rather than mining superficial cues, thus encouraging deeper reasoning and avoiding shortcut learning. Our method significantly improves multimodal reasoning accuracy and generalization, and successfully unifies multimodal fusion, cross-domain generalization, and interpretability towards trustworthy multimodal FAS. Xun Lin, Yong Xu 0001, Weicheng Xie 0001, Zitong Yu |
AAAI | 2 |
| 2026 | FaceShield: Explainable Face Anti-Spoofing with Multimodal Large Language ModelsabstractFace anti-spoofing (FAS) is crucial for protecting facial recognition systems from presentation attacks. Previous methods approached this task as a classification problem, lacking interpretability and reasoning behind the predicted results. Recently, multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and decision-making in visual tasks. However, there is currently no universal and comprehensive MLLM and dataset specifically designed for FAS task. To address this gap, we propose FaceShield, a MLLM for FAS, along with the corresponding pre-training and supervised fine-tuning (SFT) datasets, FaceShield-pre10K and FaceShield-sft45K. FaceShield is capable of determining the authenticity of faces, identifying types of spoofing attacks, providing reasoning for its judgments, and detecting attack areas. Specifically, we employ spoof-aware vision perception (SAVP) that incorporates both the original image and auxiliary information based on prior knowledge. We then use an prompt-guided vision token masking (PVTM) strategy to random mask vision tokens, thereby improving the model's generalization ability. We conducted extensive experiments on three benchmark datasets, demonstrating that FaceShield significantly outperforms previous deep learning models and general MLLMs on four FAS tasks, i.e., coarse-grained classification, fine-grained classification, reasoning, and attack localization. Hongyang Wang 0001, Zhuofu Tao, Yuhao Gao, Liepiao Zhang, Xun Lin, Xiaochen Yuan, Zitong Yu, Xiaochun Cao |
AAAI | 6 |
| 2026 | Multimodal Mixture-of-Experts with Retrieval Augmentation for Protein Active Site IdentificationabstractAccurate identification of protein active sites at the residue level is crucial for understanding protein function and advancing drug discovery. However, current methods face two critical challenges: vulnerability in single-instance prediction due to sparse training data, and inadequate modality reliability estimation that leads to performance degradation when unreliable modalities dominate fusion processes. To address these challenges, we introduce Multimodal Mixtureof-Experts with Retrieval Augmentation (MERA), the first retrieval-augmented framework for protein active site identification. MERA employs hierarchical multi-expert retrieval that dynamically aggregates contextual information from chain, sequence, and active-site perspectives through residuelevel mixture-of-experts gating. To prevent modality degradation, we propose a reliability-aware fusion strategy based on Dempster–Shafer evidence theory that quantifies modality trustworthiness through belief mass functions and learnable discounting coefficients, enabling principled multimodal integration. Extensive experiments on ProTAD-Gen and TS125 datasets demonstrate that MERA achieves state-of-the-art performance, with 90% AUPRC on active site prediction and significant gains on peptide-binding site identification, validating the effectiveness of retrieval-augmented multi-expert modeling and reliability-guided fusion Jiale Zhou 0001, Rubo Wang, Xun Lin, Tianxu Lv, Leong Hou U, Yefeng Zheng 0001 |
AAAI | 5 |
| 2026 | GoSSamer: Lightweight and Linear-Communication Asynchronous (Dynamic Proactive) Secret Sharing and the Applications
Xinxin Xing, Yizhong Liu, Boyang Liao, Jianwei Liu 0001, Bin Hu 0001, Xun Lin, Yuan Lu 0001, Tianwei Zhang 0004 |
SP | 6 |
| 2026 | ICPE-FAS: Instance and Category Prompts Engineering for Generalizable Face Anti-Spoofing
Ajian Liu 0001, Xun Lin, Hui Ma 0018, Xinxing Yu, Jiabao Guo, Zitong Yu, Jun Wan 0001, Zhanchuan Cai, Zhen Lei 0001, Yanyan Liang 0001 |
Int. J. Comput. Vis. | 2 |
| 2026 | DGPDL: Domain-Guided Prompt Distribution Learning for Generalizable Face Anti-SpoofingabstractThe overfitting of domain signals results in poor domain generalization of face anti-spoofing. The current methods usually improve the diversity of source domains to alleviate this overfitting. However, this benefit is minimal, as even the most diverse domain signals will also be absent in the target domain. In this work, we propose a Domain-Guided Prompt Distribution Learning (DGPDL) built on Vision-Language Models like CLIP, which explores a unified representation of domain signals as a prompt across the source and target domain to alleviate the understanding bias caused by domain gaps. Specifically, we first define a learnable Domain-Specific Distribution (DSD) that covers as many domain elements as possible, such as image quality, color tone, camera settings, etc., which establish connections between different domains and linearly combinable prompt in any domain; Then, based on the style statistics of the given sample, we construct its optimal Domain-Specific Prompts (DSPs) from the defined DSD through the designed Prompt Assemble Attention (PAA) with the similarity matching; Finally, the assembled DSPs will act as carrier or agent to perform on both the vision and language branches, synergistically improving the model's recognition of domain signals. By using the prompt to represent domain signals uniformly, if the model can be robust to DSPs in the source domain, it should be applicable to target domain, as they share the same DSD. By representing domain signals as prompts rather than instantiation features, DGPDL effectively reduces the reliance on specific domain appearances. This design enables the model to dynamically adapt to unseen target domains without the need for retraining. Extensive experiments show that the DGPDL is effective and outperforms the state-of-the-art methods on several cross-domain benchmarks. Ajian Liu 0001, Xun Lin, Ruicong Zhi, Yanyan Liang 0001, Xinshan Zhu, Zhanchuan Cai, Jun Wan 0001, Sergio Escalera, Zhen Lei 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Multi-Leader Byzantine Fault Tolerance in Blockchain: Performance and Security
Yizhong Liu, Mingzhe Zhai, Xun Lin, Chenhao Ying 0001, Zhenyu Guan 0002, Dawei Li 0009, Qianhong Wu, Jianwei Liu 0001, Willy Susilo, Robert H. Deng |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Backdoor Attacks Against No-Reference Image Quality Assessment Models via a Scalable TriggerabstractNo-Reference Image Quality Assessment (NR-IQA), responsible for assessing the quality of a single input image without using any reference, plays a critical role in evaluating and optimizing computer vision systems, e.g., low-light enhancement. Recent research indicates that NR-IQA models are susceptible to adversarial attacks, which can significantly alter predicted scores with visually imperceptible perturbations. Despite revealing vulnerabilities, these attack methods have limitations, including high computational demands, untargeted manipulation, limited practical utility in white-box scenarios, and reduced effectiveness in black-box scenarios. To address these challenges, we shift our focus to another significant threat and present a novel poisoning-based backdoor attack against NR-IQA (BAIQA), allowing the attacker to manipulate the IQA model's output to any desired target value by simply adjusting a scaling coefficient alpha for the trigger. We propose to inject the trigger in the discrete cosine transform (DCT) domain to improve the local invariance of the trigger for countering trigger diminishment in NR-IQA models due to widely adopted data augmentations. Furthermore, the universal adversarial perturbations (UAP) in the DCT space are designed as the trigger, to increase IQA model susceptibility to manipulation and improve attack effectiveness. In addition to the heuristic method for poison-label BAIQA (P-BAIQA), we explore the design of clean-label BAIQA (C-BAIQA), focusing on alpha sampling and image data refinement, driven by theoretical insights we reveal. Extensive experiments on diverse datasets and various NR-IQA models demonstrate the effectiveness of our attacks. Yi Yu 0011, Song Xia, Xun Lin, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
AAAI | 3 |
| 2025 | EPE-P: Evidence-based Parameter-efficient Prompting for Multimodal Learning with Missing ModalitiesabstractMissing modalities are a common challenge in real-world multimodal learning scenarios, occurring during both training and testing. Existing methods for managing missing modalities often require the design of separate prompts for each modality or missing case, leading to complex designs and a substantial increase in the number of parameters to be learned. As the number of modalities grows, these methods become increasingly inefficient due to parameter redundancy. To address these issues, we propose Evidence-based Parameter-Efficient Prompting (EPE-P), a novel and parameter-efficient method for pretrained multimodal networks. Our approach introduces a streamlined design that integrates prompting information across different modalities, reducing complexity and mitigating redundant parameters. Furthermore, we propose an Evidence-based Loss function to better handle the uncertainty associated with missing modalities, improving the model’s decision-making. Our experiments demonstrate that EPE-P outperforms existing prompting-based methods in terms of both effectiveness and efficiency. The code is released at https://github.com/Boris-Jobs/EPE-P_MLLMs-Robustness. Xun Lin, Yawen Cui, Zitong Yu |
ICASSP | 2 |
| 2025 | Big-Moe: Bypassing Isolated Gating For Generalized Multimodal Face Anti-SpoofingabstractIn the domain of facial recognition security, multimodal Face Anti-Spoofing (FAS) is essential for countering presentation attacks. However, existing technologies encounter challenges due to modality biases and imbalances, as well as domain shifts. Our research introduces a Mixture of Experts (MoE) model to address these issues effectively. We identified three limitations in traditional MoE approaches to multimodal FAS: (1) Coarse-grained experts’ inability to capture nuanced spoofing indicators; (2) Gated networks’ susceptibility to input noise affecting decision-making; (3) MoE’s sensitivity to prompt tokens leading to overfitting with conventional learning methods. To mitigate these, we propose the Bypass Isolated Gating MoE (BIG-MoE) framework, featuring: (1) Fine-grained experts for enhanced detection of subtle spoofing cues; (2) An isolation gating mechanism to counteract input noise; (3) A novel differential convolutional prompt bypass enriching the gating network with critical local features, thereby improving perceptual capabilities. Extensive experiments on four benchmark datasets demonstrate significant generalization performance improvement in multimodal FAS task. The code is released at https://github.com/murInJ/BIG-MoE. Zitong Yu, Xun Lin, Weicheng Xie 0001, LinLin Shen |
ICASSP | 3 |
| 2025 | DADM: Dual Alignment of Domain and Modality for Face Anti-SpoofingabstractWith the availability of diverse sensor modalities (i.e., RGB, Depth, Infrared) and the success of multi-modal learning, multi-modal face anti-spoofing (FAS) has emerged as a prominent research focus. The intuition behind it is that leveraging multiple modalities can uncover more intrinsic spoofing traces. However, this approach presents more risk of misalignment. We identify two main types of misalignment: (1) \textbf{Intra-domain modality misalignment}, where the importance of each modality varies across different attacks. For instance, certain modalities (e.g., Depth) may be non-defensive against specific attacks (e.g., 3D mask), indicating that each modality has unique strengths and weaknesses in countering particular attacks. Consequently, simple fusion strategies may fall short. (2) \textbf{Inter-domain modality misalignment}, where the introduction of additional modalities exacerbates domain shifts, potentially overshadowing the benefits of complementary fusion. To tackle (1), we propose a alignment module between modalities based on mutual information, which adaptively enhances favorable modalities while suppressing unfavorable ones. To address (2), we employ a dual alignment optimization method that aligns both sub-domain hyperplanes and modality angle margins, thereby mitigating domain gaps. Our method, dubbed \textbf{D}ual \textbf{A}lignment of \textbf{D}omain and \textbf{M}odality (DADM), achieves state-of-the-art performance in extensive experiments across four challenging protocols demonstrating its robustness in multi-modal domain generalization scenarios. The codes will be released soon. Xun Lin, Zitong Yu, Liepiao Zhang, Xin Liu 0012, Hui Li 0089, Xiaochen Yuan, Xiaochun Cao |
ICCV | 2 |
| 2025 | TopoTTA: Topology-Enhanced Test-Time Adaptation for Tubular Structure SegmentationabstractTubular structure segmentation (TSS) is important for various applications, such as hemodynamic analysis and route navigation. Despite significant progress in TSS, domain shifts remain a major challenge, leading to performance degradation in unseen target domains. Unlike other segmentation tasks, TSS is more sensitive to domain shifts, as changes in topological structures can compromise segmentation integrity, and variations in local features distinguishing foreground from background (e.g., texture and contrast) may further disrupt topological continuity. To address these challenges, we propose Topology-enhanced Test-Time Adaptation (TopoTTA), the first test-time adaptation framework designed specifically for TSS. TopoTTA consists of two stages: Stage 1 adapts models to cross-domain topological discrepancies using the proposed Topological Meta Difference Convolutions (TopoMDCs), which enhance topological representation without altering pre-trained parameters; Stage 2 improves topological continuity by a novel Topology Hard sample Generation (TopoHG) strategy and prediction alignment on hard samples with pseudo-labels in the generated pseudo-break regions. Extensive experiments across four scenarios and ten datasets demonstrate TopoTTA's effectiveness in handling topological distribution shifts, achieving an average improvement of 31.81% in clDice. TopoTTA also serves as a plug-and-play TTA solution for CNN-based TSS models. Jiale Zhou 0001, Wenhan Wang, Shikun Li, Xiaolei Qu, Yizhong Liu, Wenzhong Tang, Xun Lin, Yefeng Zheng 0001 |
ICCV | 8 |
| 2025 | AdaMHF: Adaptive Multimodal Hierarchical Fusion for Survival PredictionabstractThe integration of pathologic images and genomic data for survival analysis has gained increasing attention with advances in multimodal learning. However, current methods often ignore biological characteristics, such as heterogeneity and sparsity, both within and across modalities, ultimately limiting their adaptability to clinical practice. To address these challenges, we propose AdaMHF: Adaptive Multimodal Hierarchical Fusion, a framework designed for efficient, comprehensive, and tailored feature extraction and fusion. AdaMHF is specifically adapted to the uniqueness of medical data, enabling accurate predictions with minimal resource consumption, even under challenging scenarios with missing modalities. Initially, AdaMHF employs an experts expansion and residual structure to activate specialized experts for extracting heterogeneous and sparse features. Extracted tokens undergo refinement via selection and aggregation, reducing the weight of non-dominant features while preserving comprehensive information. Subsequently, the encoded features are hierarchically fused, allowing multi-grained interactions across modalities to be captured. Furthermore, we introduce a survival prediction benchmark designed to resolve scenarios with missing modalities, mirroring real-world clinical conditions. Extensive experiments on TCGA datasets demonstrate that AdaMHF surpasses current state-of-the-art (SOTA) methods, showcasing exceptional performance in both complete and incomplete modality settings. Code is available in AdaMHF. Shuaiyu Zhang, Xun Lin, Rongxiang Zhang, Yong Xu 0001, Tao Tan 0002, Xubin Zheng, Zitong Yu |
ICME | 2 |
| 2025 | Dynamic Analysis and Adaptive Discriminator for Fake News DetectionabstractIn current web environment, fake news spreads rapidly across online social networks, posing serious threats to society. Existing multimodal fake news detection methods can generally be classified into knowledge-based and semantic-based approaches. However, these methods are heavily rely on human expertise and feedback, lacking flexibility. To address this challenge, we propose a Dynamic Analysis and Adaptive Discriminator (DAAD) approach for fake news detection. For knowledge-based methods, we introduce the Monte Carlo Tree Search algorithm to leverage the self-reflective capabilities of large language models (LLMs) for prompt optimization, providing richer, domain-specific details and guidance to the LLMs, while enabling more flexible integration of LLM comment on news content. For semantic-based methods, we define four typical deceit patterns: emotional exaggeration, logical inconsistency, image manipulation, and semantic inconsistency, to reveal the mechanisms behind fake news creation. To detect these patterns, we carefully design four discriminators and expand them in depth and breadth, using the soft-routing mechanism to explore optimal detection models. Experimental results on three real-world datasets demonstrate the superiority of our approach. Xinqi Su, Zitong Yu, Yawen Cui, Ajian Liu 0001, Xun Lin, Haochen Liang, Wenhui Li 0001, Li Shen 0008, Xiaochun Cao |
ACM Multimedia | 5 |
| 2025 | Fully Anonymous Decentralized Identity Supporting Threshold Traceability with Practical Blockchain
Yizhong Liu, Zedan Zhao, Feiang Ran, Xun Lin, Dawei Li 0009, Zhenyu Guan 0002 |
WWW | 5 |
| 2025 | Quantum-Resistant Sharding Blockchain and Its Application in Secure Data TransmissionabstractWith the approach of the quantum era, public key cryptography (PKC) faces risks, which also presents challenges to blockchain technologies that utilize PKC as a core component. Sharding blockchain is a promising way to realize scalability, yet current research does not consider quantum-resistant sharding blockchains as it is non-trivial to design cross-shard communication and transaction processing method without PKC. Besides, blockchain enables reliability in data transmission and unbreakable communication while current schemes suffer from high overhead and low throughput. In this paper, we propose a quantum-resistant sharding blockchain (QRShar) and a secure data transmission scheme (QRDT) to fill the above gap. Firstly, we design a secure and efficient cross-shard communication pattern utilizing hash-based message authentication code (HMAC) and erasure code to reduce the transmission load and achieve high efficiency. Secondly, we propose the a quantum-resistant sharding blockchain utilizing optimized cross-shard transaction processing method to decrease the consensus execution frequency. Thirdly, we introduce a quantum-resistant key agreement protocol through the verifiable secret sharing on cryptographic hash function and we also offer a data transmission scheme to realize efficient QRDT. Furthermore, we conduct security analysis and performance evaluations for our schemes. The results show that the QRShar throughput can reach up to 34 KTPS and the latency stays below 2 seconds. The key agreement latency is just 43ms. Yizhong Liu, Xun Lin, Zhenyu Guan 0002, Dawei Li 0009, Jianwei Liu 0001, Qianhong Wu, Willy Susilo, Robert H. Deng |
IEEE J. Sel. Areas Commun. | 3 |
| 2025 | Rehearsal-Free and Efficient Continual Learning for Cross-Domain Face Anti-SpoofingabstractFace Anti-Spoofing (FAS) is constantly challenged by new attack types and mediums, and thus it is crucial for a FAS model to not only mitigate Catastrophic Forgetting (CF) of previously learned spoofing knowledge on the training data during continual learning but also enhance the model's generalization ability to potential spoofing attacks. In this paper, we first highlight that current strategies for catastrophic forgetting are not well-suited to the imperceptible nature of spoofing information in FAS and lack the focus on improving generalization capability. Then, the instance-wise dynamic central difference convolutional adapter module with the weighted ensemble strategy for Vision Transformer (ViT) is proposed for efficiently fine-tuning with low-shot data by extracting generalized spoofing texture information. Furthermore, we find that catastrophic forgetting in FAS can be reflected through the inconsistent attention matrices of ViT between different continual sessions, as the attention matrices embody relationships of spoofing clues between different patch tokens. Hence, we introduce attention consistency regularization by learning and reusing attention matrices to alleviate catastrophic forgetting. Finally, we devise new protocols and conduct extensive experiments to validate the superior performance of alleviating catastrophic forgetting and generalization on unseen domains. Rizhao Cai, Yawen Cui, Zitong Yu, Xun Lin, Changsheng Chen 0001, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Reliable and Balanced Transfer Learning for Generalized Multimodal Face Anti-SpoofingabstractFace Anti-Spoofing (FAS) is essential for securing face recognition systems against presentation attacks. Recent advances in sensor technology and multimodal learning have enabled the development of multimodal FAS systems. However, existing methods often struggle to generalize to unseen attacks and diverse environments due to two key challenges: (1) Modality unreliability, where sensors such as depth and infrared suffer from severe domain shifts, impairing the reliability of cross-modal fusion; and (2) Modality imbalance, where over-reliance on a dominant modality weakens the model's robustness against attacks that affect other modalities. To overcome these issues, we propose MMDG++, a multimodal domain-generalized FAS framework built upon the vision-language model CLIP. In MMDG++, we design the Uncertainty-Guided Cross-Adapter++ (U-Adapter++) to filter out unreliable regions within each modality, enabling more reliable multimodal interactions. Additionally, we introduce Rebalanced Modality Gradient Modulation (ReGrad) for adaptive gradient modulation to balance modality convergence. To further enhance generalization, propose Asymmetric Domain Prompts (ADPs) that leverage CLIP's language priors to learn generalized decision boundaries across modalities. We also develop a novel multimodal FAS benchmark to evaluate generalizability under various deployment conditions. Extensive experiments across this benchmark show our method outperforms state-of-the-art FAS methods, demonstrating superior generalization capability. Xun Lin, Ajian Liu 0001, Zitong Yu, Rizhao Cai, Shuai Wang 0049, Yi Yu 0011, Jun Wan 0001, Zhen Lei 0001, Xiaochun Cao, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Toward Model Resistant to Transferable Adversarial Examples via Trigger ActivationabstractAdversarial examples, characterized by imperceptible perturbations, pose significant threats to deep neural networks by misleading their predictions. A critical aspect of these examples is their transferability, allowing them to deceive unseen models in closed-box scenarios. Despite the widespread exploration of defense methods, including those on transferability, they show limitations: inefficient deployment, ineffective defense, and degraded performance on clean images. In this work, we introduce a novel training paradigm aimed at enhancing robustness against transferable adversarial examples (TAEs) in a more efficient and effective way. We propose a model that exhibits random guessing behavior when presented with clean data$\boldsymbol {x}$as input, and generates accurate predictions when with triggered data$\boldsymbol {x}+\boldsymbol {\tau }$. Importantly, the trigger$\boldsymbol {\tau }$remains constant for all data instances. We refer to these models as models with trigger activation. We are surprised to find that these models exhibit certain robustness against TAEs. Through the consideration of first-order gradients, we provide a theoretical analysis of this robustness. Moreover, through the joint optimization of the learnable trigger and the model, we achieve improved robustness to transferable attacks. Extensive experiments conducted across diverse datasets, evaluating a variety of attacking methods, underscore the effectiveness and superiority of our approach. Yi Yu 0011, Song Xia, Xun Lin, Chenqi Kong, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | MorFormer: Morphology-Aware Transformer for Generalized Pavement Crack SegmentationabstractCracks are common on pavements. Accurate crack detection plays a vital role in pavement maintenance. However, cracks have rich and varied morphological features and fine edges, making this task challenging. Additionally, noise factors such as stains, scratches, and complex textures in the pavement background can easily be confused with cracks, increasing the risk of false prediction in the segmentation process. Therefore, we propose Background Morphology Learning (BML) to reconstruct morphological features of the pavement background noise, extract background morphological dissimilarity maps to suppress interference and reduce false alarms. In addition, we propose Crack Morphology-aware Attention (CMA), which adaptively learns the morphological shape of cracks and dynamically adjusts the shape of the attention receptive field to the topological features of the cracks. This significantly improves the completeness of segmentation. Our method mitigates the problems of false alarms and incomplete segmentation results in the crack segmentation task. Therefore, we propose a Morphology-Aware Transformer (MorFormer) that achieves state-of-the-art results on five public datasets. Moreover, we propose a large-scale cross-domain benchmark for crack segmentation, where MorFormer exhibits excellent domain generalization. Wenzhong Tang, Shuai Wang 0049, Xiaolei Qu, Xun Lin |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | Suppress and Rebalance: Towards Generalized Multi-Modal Face Anti-SpoofingabstractFace Anti-Spoofing (FAS) is crucial for securing face recognition systems against presentation attacks. With ad-vancements in sensor manufacture and multi-modal learning techniques, many multi-modal FAS approaches have emerged. However, they face challenges in generalizing to unseen attacks and deployment conditions. These chal-lenges arise from (1) modality unreliability, where some modality sensors like depth and infrared undergo signifi-cant domain shifts in varying environments, leading to the spread of unreliable information during cross-modal feature fusion, and (2) modality imbalance, where training overly relies on a dominant modality hinders the conver-gence of others, reducing effectiveness against attack types that are indistinguishable by sorely using the dominant modality. To address modality unreliability, we propose the Uncertainty-Guided Cross-Adapter (U-Adapter) to recognize unreliably detected regions within each modality and suppress the impact of unreliable regions on other modal-ities. For modality imbalance, we propose a Rebalanced Modality Gradient Modulation (ReGrad) strategy to rebal-ance the convergence speed of all modalities by adaptively adjusting their gradients. Besides, we provide the first large-scale benchmark for evaluating multi-modal FAS per-formance under domain generalization scenarios. Exten-sive experiments demonstrate that our method outperforms state-of-the-art methods. Source codes and protocols are released on https://github.com/OMGGGGG/mmdg. Xun Lin, Shuai Wang 0049, Rizhao Cai, Yizhong Liu, Ying Fu 0001, Wenzhong Tang, Zitong Yu, Alex Chichung Kot |
CVPR | 1 |
| 2024 | Flexible-Modal Deception Detection with Audio-Visual AdapterabstractDeception detection within audio-visual modalities is vital across diverse sectors, notably in customs security and multimedia anti-fraud. However, this notable efficacy is lost by the necessity to train and deploy separate models for each conceivable modality scenario, leading to redundancy and inefficiency. Moreover, real-world environments where multi-modal models are deployed often fail to meet these idealized conditions. To overcome these challenges and further elevate performance levels, we propose an advanced Transformer-based framework complemented by an Audio-Visual Adapter (AVA) integrating temporal features from both audio and visual modalities. In addition, we introduce an innovative multi-modal contrastive learning method that is designed to enhance the correlation between uni-modal features and their integrated counterparts within a consistent feature space. Our designed method can deal with the flexible-model scenario instead of deploying different models for various modalities. Empirical evaluations conducted on two benchmark datasets have validated the superiority of our proposed model over other multi-modal fusion techniques, particularly in scenarios characterized by varying and missing modalities. This strongly affirms the effectiveness of our approach in significantly boosting the accuracy of deception detection in complex, real-world multi-modal scenarios. The codes will be released soon. Zhaoxu Li, Zitong Yu, Xun Lin, Nithish Muthuchamy Selvaraj, Xiaobao Guo, Bingquan Shen, Adams Wai-Kin Kong, Alex Chichung Kot |
IJCB | 3 |
| 2024 | DDAP: Dual-Domain Anti-Personalization against Text-to-Image Diffusion ModelsabstractDiffusion-based personalized visual content generation technologies have achieved significant breakthroughs, allowing for the creation of specific objects by just learning from a few reference photos. However, when misused to fabricate fake news or unsettling content targeting individuals, these technologies could cause considerable societal harm. To address this problem, current methods generate adversarial samples by adversarially maximizing the training loss, thereby disrupting the output of any personalized generation model trained with these samples. However, the existing methods fail to achieve effective defense and maintain stealthiness, as they overlook the intrinsic properties of diffusion models. In this paper, we introduce a novel Dual-Domain Anti-Personalization framework (DDAP). Specifically, we have developed Spatial Perturbation Learning (SPL) by exploiting the fixed and perturbation-sensitive nature of the image encoder in personalized generation. Subsequently, we have designed a Frequency Perturbation Learning (FPL) method that utilizes the characteristics of diffusion models in the frequency domain. The SPL disrupts the overall texture of the generated images, while the FPL focuses on image details. By alternating between these two methods, we construct the DDAP framework, effectively harnessing the strengths of both domains. To further enhance the visual quality of the adversarial samples, we design a localization module to accurately capture attentive areas while ensuring the effectiveness of the attack and avoiding unnecessary disturbances in the background. Extensive experiments on facial benchmarks have shown that the proposed DDAP enhances the disruption of personalized generation models while also maintaining high quality in adversarial samples, making it more effective in protecting privacy in practical applications. Runping Xi, Yingxin Lai, Xun Lin, Zitong Yu |
IJCB | 4 |
| 2024 | CPL-CLIP: Compound Prompt Learning for Flexible-Modal Face Anti-SpoofingabstractFace anti-spoofing (FAS) is pivotal in safeguarding the integrity of face recognition systems. Flexible-modal FAS utilizes multi-modal data and trains a unified model adaptable to any single-modal testing scenario. This innovation addresses the shortcomings of conventional multi-modal FAS approaches, which typically demand separate model training and deployment for each modality. However, existing flexible-modal FAS approaches activate specific network branches based on the modality of the tested sample. This not only increases the model’s parameters but also necessitates the provision of the image’s modality for testing, thereby constraining deployment flexibility. To address the issue, we present Compound Prompt Learning CLIP (CPL-CLIP), a novel method for flexible-modal FAS. This approach capitalizes on a learned textual prompt that is nearly independent of modality, thus bolstering class-based classification across arbitrary modalities. Specifically, our CPL-CLIP introduces a Dual-Branch Prompt (DBP), consisting of class and modal prompts that describe and guide classification, where each prompt is composed of learnable vectors and fixed templates. To further render the class prompt as modality-agnostic as possible, a Cosine Similarity Loss (CSL) is proposed to facilitate the maximal separation of the class prompt from the modality prompt. With only the class prompt utilized during testing, CPL-CLIP enables deployment in diverse modal testing scenarios without the necessity of the test image’s modality to be known. Extensive experiments demonstrate CPL-CLIP’s superiority over existing methods on several flexible-modal FAS benchmarks. Xiangyu Zhu 0001, Ajian Liu 0001, Xun Lin, Jun Wan 0001, Zhen Lei 0001 |
IJCB | 4 |
| 2024 | HideMIA: Hidden Wavelet Mining for Privacy-Enhancing Medical Image Analysis
Xun Lin, Yi Yu 0011, Zitong Yu, Ruohan Meng, Jiale Zhou 0001, Ajian Liu 0001, Yizhong Liu, Shuai Wang 0049, Wenzhong Tang, Zhen Lei 0001, Alex Chichung Kot |
ACM Multimedia | 1 |
| 2024 | Transferable Adversarial Attacks on SAM and Its Downstream ModelsabstractThe utilization of large foundational models has a dilemma: while fine-tuning downstream tasks from them holds promise for making use of the well-generalized knowledge in practical applications, their open accessibility also poses threats of adverse usage.
This paper, for the first time, explores the feasibility of adversarial attacking various downstream models fine-tuned from the segment anything model (SAM), by solely utilizing the information from the open-sourced SAM.
In contrast to prevailing transfer-based adversarial attacks, we demonstrate the existence of adversarial dangers even without accessing the downstream task and dataset to train a similar surrogate model.
To enhance the effectiveness of the adversarial attack towards models fine-tuned on unknown datasets, we propose a universal meta-initialization (UMI) algorithm to extract the intrinsic vulnerability inherent in the foundation model, which is then utilized as the prior knowledge to guide the generation of adversarial perturbations.
Moreover, by formulating the gradient difference in the attacking process between the open-sourced SAM and its fine-tuned downstream models, we theoretically demonstrate that a deviation occurs in the adversarial update direction by directly maximizing the distance of encoded feature embeddings in the open-sourced SAM.
Consequently, we propose a gradient robust loss that simulates the associated uncertainty with gradient-based noise augmentation to enhance the robustness of generated adversarial examples (AEs) towards this deviation, thus improving the transferability.
Extensive experiments demonstrate the effectiveness of the proposed universal meta-initialized and gradient robust adversarial attack (UMI-GRAT) toward SAMs and their downstream models.
Code is available at https://github.com/xiasong0501/GRAT. Song Xia, Wenhan Yang, Yi Yu 0011, Xun Lin, Henghui Ding, Ling-Yu Duan, Xudong Jiang 0001 |
NeurIPS | 4 |
| 2024 | A MLP architecture fusing RGB and CASSI for computational spectral imaging
Zeyu Cai 0001, Ru Hong, Xun Lin, Jiming Yang, Youliang Ni, Chengqian Jin, Feipeng Da |
Comput. Vis. Image Underst. | 3 |
| 2024 | SS-DID: A Secure and Scalable Web3 Decentralized Identity Utilizing Multilayer Sharding BlockchainabstractWeb3 is a revolutionary Internet paradigm that focusing decentralization, user empowerment, and intelligence. One of its key technologies is decentralized identity (DID), which has gained significant attention recently. However, existing DID solutions are not scalable enough to be compatible with the large-scale identity node applications required by Web3 across various fields. To overcome this challenge, we propose the first multi-layer Web3 DID architecture utilizing sharding blockchain, which provides management, scalability, and compatibility. This architecture leverages leader shards and the main chain to establish trust, while regular shards manage DID-related transactions. Specific system processes and query optimizations are also given. Besides, formal security analysis and comprehensive simulation evaluations have demonstrated that the architecture can achieve all proposed security and performance goals, including low latency of down to 2 seconds and high throughput of up to 90KTPS. Yizhong Liu, Zedan Zhao, Jianwei Liu 0001, Xun Lin, Qianhong Wu, Willy Susilo |
IEEE Internet Things J. | 5 |
| 2024 | BFL-SA: Blockchain-based federated learning via enhanced secure aggregation
Yizhong Liu, Zixiao Jia, Zixu Jiang, Xun Lin, Jianwei Liu 0001, Qianhong Wu, Willy Susilo |
J. Syst. Archit. | 4 |
| 2024 | Secure and Scalable Cross-Domain Data Sharing in Zero-Trust Cloud-Edge-End Environment Based on Sharding BlockchainabstractThe cloud-edge-end architecture is suitable for many essential scenarios, such as 5 G, the Internet of Things (IoT), and mobile edge computing. Under this architecture, cross-domain and cross-layer data sharing is commonly in need. Considering cross-domain data sharing under the zero-trust model, where each entity does not trust the others, existing solutions have certain problems regarding security, fairness, scalability, and efficiency. Aiming at solving these issues, we conduct the following research. First, a new plaintext checkable encryption scheme is constructed, which can be used on lightweight IoT devices to verify the ciphertext validity sent by a data owner. Second, we propose a new multi-domain cloud-edge-end architecture based on sharding blockchains and design a cross-domain data sharing scheme under the partial trust model to achieve security, scalability, and high performance. Third, a cross-domain data sharing scheme under the zero trust model is further designed, which can ensure the fairness of both parties in data sharing. Fourth, we give a formal security definition and analysis of cross-domain data sharing. Fifth, we conduct a detailed theoretical analysis of the protocol and give an in-depth functional test and performance test, including the throughput and latency of data sharing policy registration and execution. Yizhong Liu, Xinxin Xing, Ziheng Tong, Xun Lin, Jing Chen 0003, Zhenyu Guan 0002, Qianhong Wu, Willy Susilo |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2023 | Image manipulation detection by multiple tampering traces and edge artifact enhancement
Xun Lin, Shuai Wang 0049, Jiahao Deng, Ying Fu 0001, Xiao Bai 0001, Xinlei Chen, Xiaolei Qu, Wenzhong Tang |
Pattern Recognit. | 1 |
| 2023 | CDS-Net: Cooperative dual-stream network for image manipulation detection
Jiahao Deng, Xun Lin, Wenzhong Tang, Shuai Wang 0049 |
Pattern Recognit. Lett. | 3 |
| 2020 | Social car: The research of interaction design on the driver's communication systemabstractSummary Road traffic in the city is changing with each passing day. The current communication system is based on sound, light, and other ways. Insufficient information may lead to misunderstandings and conflicts among drivers. We hope to use the existing digital information and interaction technology in designing socialized information service for automobiles. This is to enable drivers to exercise a more social behavior while driving. In this paper, the interaction design of a driver communication system is studied. We carried out three driving simulation experiments to study three different ways of interaction interfaces and then promoted a questionnaire based on NASA TLX to understand the subjective feelings and subjective workload of drivers in using these interaction interfaces. The four dimensions of driving quality evaluation, road attention evaluation, interaction effect evaluation, and subjective evaluation are analyzed. The results suggest four ideas. (a) The distractions of the audio and text interface are far below that of the video interface. The audio interface has the least distraction by sending information. (b) The errors in words or pronunciation barely affect the understanding of the meaning of the drivers' message. However, it will cause an increase in distraction. (c) Marking some important information in text display gives drivers a better understanding and interaction experience. (d) The video interface gives drivers a sense of social participation. According to these conclusions, we can give some guidance to the driver's communication system. It not only helps the drivers accomplish the driving task but also creates more social actions with other drivers. It is necessary to reduce driving distraction and enhance the availability for the communication system in the future. Xun Lin |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | A Distributed Approach Toward Discriminative Distance Metric LearningabstractDistance metric learning (DML) is successful in discovering intrinsic relations in data. However, most algorithms are computationally demanding when the problem size becomes large. In this paper, we propose a discriminative metric learning algorithm, develop a distributed scheme learning metrics on moderate-sized subsets of data, and aggregate the results into a global solution. The technique leverages the power of parallel computation. The algorithm of the aggregated DML (ADML) scales well with the data size and can be controlled by the partition. We theoretically analyze and provide bounds for the error induced by the distributed treatment. We have conducted experimental evaluation of the ADML, both on specially designed tests and on practical image annotation tasks. Those tests have shown that the ADML achieves the state-of-the-art performance at only a fraction of the cost incurred by most existing methods. Jun Li 0010, Xun Lin, Xiaoguang Rui, Yong Rui, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |