Shuai Liu 0016

dblp:76/5789-16 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-0327-6729ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Security and privacy · 5 · 1 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 EchoBat: Echo-Vision Enhancement and Echo-Layered Sampling for Video LLMs Hallucination Mitigation
abstract
Recent advancements in multimodal large language models (MLLMs) have shown remarkable progress in video understanding. However, video MLLMs (VideoMLLMs) still suffer from hallucinations, generating nonsensical or irrelevant content. This issue partly stems from over-reliance on pre-trained knowledge, sometimes neglecting the rich visual information present in the video. Additionally, many existing methods rely on uniform frame sampling, which can overlook critical visual cues. To address these challenges, we present EchoBat, a novel approach that leverages audio information as well as video temporal and logical consistency to improve preference data construction and keyframe extraction. Our method integrates Direct Preference Optimization (DPO) to mitigate hallucinations by leveraging high-quality, contextually rich preference feedback. Specifically, we use GPT-4o to generate high-quality video descriptions and integrate visually relevant segments from Whisper-derived transcripts to construct preference responses. Correspondingly, we use the reference model itself to describe the reversed video, and use GPT-4o to flashback the text and fill in the hallucination to produce non-preferred responses. This strategy enhances the model’s ability to better understand visual content and temporal, logical relationships within videos. Furthermore, we propose an echo-layered sampling strategy for keyframe extraction from videos, which can provide more precise visual supervision compared to uniform sampling. Experimental results on the three latest video hallucination benchmarks demonstrate the effectiveness of our approach.
Shuai Liu 0016, Yiheng Pan, Chenwei Tian, Qian Li 0024, Chenhao Lin
AAAI1
2026 CLIP-ADA: CLIP-Guided Artifact-Invariant Generalizable Synthetic Image Detection
abstract
The rapid advancement of generative models necessitates detection methods that generalize to synthetic images containing diverse generator and semantic artifacts. Recent research has leveraged pre-trained vision-language models, such as CLIP, to extract forensic features that distinguish real and fake images, illustrating their promising performance in synthetic image detection. However, a systematic investigation into the embedding space of CLIP to guide its principled utilization for synthetic image detection remains largely unexplored. This paper addresses this gap by first analyzing the multi-stage CLIP image embedding space to uncover its relationship with cross-artifact forensic patterns. Our findings reveal that the mid-level stages primarily encode forensic and generator artifact features, while the high-level stages primarily encode semantic artifact features. Building upon these insights, we propose the CLIP-guided Dual-level Augmentation and Forensic Distribution Adaptation (CLIP-ADA) framework to perform artifact-invariant generalizable detection. Specifically, dual-level augmentation diversifies fake embeddings and suppresses artifact encoding during training to mitigate detectors from excessively relying on artifact features. Moreover, forensic distribution adaptation reformulates synthetic image detection as identifying distributional deviations from the CLIP encoded real embeddings and thereby designing adapters to extract cross-artifact forensic features in a detection scenario-adaptive manner. Extensive evaluations on both the conventional single-generator and continual learning-based multi-generator training settings demonstrate the effectiveness of our method, both suppressing the state-of-the-art methods by over 6% of average accuracy on unseen data from more than 10 generators.
Jingyi Deng, Chenken Xu, Chenhao Lin, Zhengyu Zhao 0001, Shuai Liu 0016, Qian Wang 0002, Chao Shen 0001
IEEE Trans. Inf. Forensics Secur.6
2026 Cross-Region Feature Reformer With Semantic Preservation for Adversarial Malware Detection
Qian Li 0024, Di Wu 0062, Chenhao Lin, Shuai Liu 0016, Cong Wang 0001, Chao Shen 0001
IEEE Trans. Inf. Forensics Secur.4
2026 Vul-CTG: A Multimodal Framework for Software Vulnerability Detection via Code Text and Graph Integration
abstract
Pretrained Language Models (PLMs) and Graph Neural Networks (GNNs) have emerged as promising approaches for software vulnerability detection. However, existing methods still face limitations, including the absence of fine-grained cross-modal interaction and the impact of data noise. Approaches integrating PLMs and GNNs fail to fully leverage their complementary strengths, while unreliable labels hinder generalization, further degrading real-world detection performance. To over-come these limitations, we propose Vul-CTG, a multimodal integration framework for software vulnerability detection that combines Code Text, and program Graph representations. Vul-CTG constructs enriched code graph representations by integrating statement-level source code graphs and abstract code property graphs, enabling more effective alignment between structural and semantic information. To enhance robustness against noisy labels and improve cross-modal consistency, the model incorporates contrastive learning and pre-training techniques. Central to Vul-CTG is CTG-Former, a novel alignment architecture that projects both code text and graph modalities into a unified latent space, allowing the model to capture complex structural and semantic patterns for more accurate vulnerability detection. Experimental results on recent function-level datasets demonstrate the effectiveness of Vul-CTG, showing an approximate 3% improvement in F1-score over state-of-the-art methods. Our code is available at https://github.com/ryxFry/Vul-CTG.
Shuai Liu 0016, Qian Li 0024, Xinlei He 0001, Xiaoyu Zhang 0013, Chenhao Lin, Chao Shen 0001
IEEE Trans. Inf. Forensics Secur.1
2026 Adversarial Video Promotion Against Text-to-Video Retrieval
Qiwei Tian, Chenhao Lin, Zhengyu Zhao 0001, Shuai Liu 0016, Qian Li 0024, Chao Shen 0001
IEEE Trans. Inf. Forensics Secur.4
2026 Graph Attention Network-Driven Hierarchical Learning for Anti-Jamming UAV Communications
abstract
Jamming attacks pose a significant threat to the security of air-ground communications, where the challenge becomes more severe when involving multiple unmanned aerial vehicles (UAVs) incurring complex interference. To address this issue, this paper proposes a graph attention-based reinforcement learning strategy for anti-jamming UAV communications. Specifically, we consider the multi-UAV transmission and deployment in the presence of jamming attacks. Then, we formulate a zero-sum game with the legitimate side and adversary to maximize and minimize the overall transmission rate, respectively. Given the complicated structure of the game, we decompose it into two layers, tackled in a hierarchical learning framework. Particularly, the inner layer addresses the legitimate beamforming, for which we establish the graph attention network (GAT) to track the complicated interference and jamming relationship based on the graph representation of the UAV network. The outer layer address the legitimate UAV deployment and adversarial jamming policy, which is reinterpreted in a multi-agent deep reinforcement learning framework to obtain the strategies of both sides. The inner GAT is then nested within the outer multi-agent learning framework in a hierarchical manner to approximate the equilibrium of the original game model. Simulation results demonstrate the convergence and the performance superiority of the proposed learning scheme in terms of anti-jamming transmission rate. Also, the results exhibit significant generalization capability to cover different network configurations and parameters with reliable communication performance.
Xiao Tang 0001, Chao Shen 0001, Chenhao Lin, Shuai Liu 0016, Bohui Wang, Dusit Niyato, Zhu Han 0001
IEEE Trans. Wirel. Commun.5
2025 TGDrag: Adding Semantic Control into Point-based Image Editing via Text Guidance
abstract
Controllable image generation has emerged as a cutting-edge subject of interest. Current interactive point-based image editing frameworks, such as DragGAN, achieve impressive results in fine-grained and controllable image editing. However, relying solely on point-based manipulations can lead to unintended outcomes due to the inherent lack of the users’ semantic intent. To address this issue, we introduce Text-Guided Drag (TGDrag), a novel approach to adding semantic control into point-based image editing by using text prompts to guide the manipulation of handle and target points. Specifically, we design a channel correlation calculator that adaptively selects channels for the text and points to mitigate the potential influence of semantic control on point control. Furthermore, we introduce a text loss function to minimize the discrepancy between the generated images and the text prompts. Experimental results demonstrate that TGDrag achieves the expected function of semantic control while maintaining effectiveness regarding point control.
Chenhao Lin, Yanjie Zhu, Yingmao Miao, Zhengyu Zhao 0001, Shuai Liu 0016, Chao Shen 0001
ICASSP5
2025 One-Shot Face Avatar Generation in a Single Forward Pass with Identity Preservation
abstract
Face avatar generation has gained significant attention recently. With the help of the Neural Radiance Field (NeRF), existing 3D methods alleviate facial distortion in 2D methods under large pose changes. However, the state-of-the-art 3D methods still require additional optimization for generation on each given portrait, even in a one-shot manner. To address this research gap, we propose a novel one-shot approach, which achieves effective face avatar generation in only a single forward pass. This is made possible by introducing an inversion encoder trained on a large-scale dataset for accurate latent code estimation and an expression animator for accurate expression control. Our approach is also designed for better preservation of the face identity by training an additional 3D feature refiner based on cross-attention. Experimental results demonstrate the superiority of our approach in terms of 3D consistency, identity similarity, and image quality.
Yingmao Miao, Chenhao Lin, Zhengyu Zhao 0001, Shuai Liu 0016, Chao Shen 0001, Xiaohong Guan
ICASSP5
2025 CountSE: Soft Exemplar Open-Set Object Counting
Shuai Liu 0016, Shiwei Zhang 0004, Wei Ke 0003
ICCV1
2025 D3: Training-Free AI-Generated Video Detection Using Second-Order Features
Chende Zheng, Ruiqi Suo, Chenhao Lin, Zhengyu Zhao 0001, Le Yang 0007, Shuai Liu 0016, Cong Wang 0001, Chao Shen 0001
ICCV6
2025 Backdoor threats in large language models - a survey
Shuai Liu 0016, Yiheng Pan, Kun Hong, Ruite Fei, Chenhao Lin, Qian Li 0024, Chao Shen 0001
Sci. China Inf. Sci.1
2025 Robust Adversarial Defenses in Federated Learning: Exploring the Impact of Data Heterogeneity
abstract
Federated Learning (FL) enables geographically distributed clients to collaboratively train machine learning models by exchanging local model parameters while preserving data privacy. In practice, FL faces two critical challenges. First, it is vulnerable to security issues as malicious clients would artificially harm the functionality of FL by launching poisoning attacks. Second, the inherent data heterogeneity among clients (termed Non-IID data in FL) naturally arises from distributed data ownership and significantly degrades model convergence and accuracy. However, with studies separately devoted to these two research lines, the interplay between data heterogeneity and security remains poorly understood. In this paper, we systematically investigate the relationship between data heterogeneity and adversarial robustness in FL. Specifically, we propose novel data partitioning algorithms that simulate Label-Conditional Non-IID and Feature-Conditional Non-IID with quantifiable heterogeneity levels. Further, we conduct extensive experiments to evaluate classical defense methods in the practical FL environment under state-of-the-art untargeted attacks. With results in various settings, we separately analyze the connection between Non-IID to defenses and attacks. Regarding attacks, with similar effects on models, Non-IID impacts the training in a different way compared with attacks. The interaction between attacks and Non-IID provides an opportunity to cause severe damage to FL. Regarding defenses, Non-IID induces heterogeneity in model distribution among clients which raises the difficulty of maintaining fidelity and robustness for defense methods.
Qian Li 0024, Di Wu 0062, Dawei Zhou 0004, Chenhao Lin, Shuai Liu 0016, Cong Wang 0001, Chao Shen 0001
IEEE Trans. Inf. Forensics Secur.5
2024 Boosting Semi-supervised Crowd Counting with Scale-based Active Learning
abstract
The core of active semi-supervised crowd counting is the sample selection criteria. However, the scale factor has been neglected in active learning approaches despite the fact that the scale of heads varies drastically in the crowd images. In this paper, we propose a simple yet effective active labeling strategy to explicitly select informative unlabeled images, guided by the intra-scale uncertainty and inter-scale inconsistency metrics. The intra-scale uncertainty is quantified through the sum of the query-level entropy of images at different scales. Images are initially ranked based on this uncertainty for preselection. Inter-scale inconsistency is measured by the divergence between the query-level predictions of upscaled and downscaled images, allowing for the identification of the most informative images exhibiting the highest inconsistency. Additionally, we implement a progressive updating scheme for the semi-supervised crowd counting framework, in which the pseudo-labels for unlabeled images are refined iteratively. It further improves the counting accuracy. Through extensive experiments on widely used benchmarks, the proposed approach has demonstrated superior performance compared to previous state-of-the-art semi-supervised and active semi-supervised crowd counting methods.
Shiwei Zhang 0004, Wei Ke 0003, Shuai Liu 0016, Xiaopeng Hong, Tong Zhang 0023
ACM Multimedia3
2024 Breaking Semantic Artifacts for Generalized AI-generated Image Detection
abstract
With the continuous evolution of AI-generated images, the generalized detection of them has become a crucial aspect of AI security. Existing detectors have focused on cross-generator generalization, while it remains unexplored whether these detectors can generalize across different image scenes, e.g., images from different datasets with different semantics. In this paper, we reveal that existing detectors suffer from substantial Accuracy drops in such cross-scene generalization. In particular, we attribute their failures to ''semantic artifacts'' in both real and generated images, to which detectors may overfit. To break such ''semantic artifacts'', we propose a simple yet effective approach based on conducting an image patch shuffle and then training an end-to-end patch-based classifier. We conduct a comprehensive open-world evaluation on 31 test sets, covering 7 Generative Adversarial Networks, 18 (variants of) Diffusion Models, and another 6 CNN-based generative models. The results demonstrate that our approach outperforms previous approaches by 2.08\% (absolute) on average regarding cross-scene detection Accuracy. We also notice the superiority of our approach in open-world generalization, with an average Accuracy improvement of 10.59\% (absolute) across all test sets. Our code is available at *https://github.com/Zig-HS/FakeImageDetection*.
Chende Zheng, Chenhao Lin, Zhengyu Zhao 0001, Shuai Liu 0016, Chao Shen 0001
NeurIPS6
2021 Dual-graph convolutional network based on band attention and sparse constraint for hyperspectral band selection
Jie Feng 0003, Zhanwei Ye, Shuai Liu 0016, Xiangrong Zhang, Jiantong Chen, Ronghua Shang, Licheng Jiao
Knowl. Based Syst.3
2017 Semi-supervised double sparse graphs based discriminant analysis for dimensionality reduction
Puhua Chen, Licheng Jiao, Fang Liu 0001, Jiaqi Zhao 0001, Shuai Liu 0016
Pattern Recognit.6