EDBT 2026 Demo / reviewers in the wild / expert
Shuai Xiao 0001
dblp:120/4356-1
· DBLP profile ↗
33ranked-venue papers
4as first author
33since 2021 · last 2026
0000-0003-4058-8120ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Computer networks · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automatic Translational Correction of Multi-View Coronary Angiography Based on Auto-Annotation Data GenerationabstractMulti-view automatic translational correction (ATC) in coronary angiography (CAG) is critical for intraoperative automatic diagnosis, in which deep learning playing a key role. However, heartbeat-induced soft matching errors and costly annotations make it difficult to build high-quality, large-scale datasets for calibration algorithm training. The training of clinical models is difficult to fulfill, as existing datasets differ significantly from real CAG in both style and structure. To address this challenge, we propose a novel high-quality data synthesis method for annotation-free ATC. We fully automated the construction of a labeled, high-fidelity dataset for training matching models. An evolutionary algorithm is introduced for global optimization of translation estimation, mitigating epipolar constraint violations caused by vascular deformation and enabling reliable correction across large viewpoint differences. Furthermore, a theoretical analysis is presented, demonstrating that error propagation between adjacent views is more accurate than direct estimation across distant views. Our experiments on clinical datasets demonstrate that our method not only significantly outperforms weakly supervised learning approaches, but also performs comparably to fully supervised methods. Moreover, it exhibits remarkable multicenter generalizability. Zhuo Zhang 0025, Shuai Xiao 0001, Jialin Li 0002, Guipeng Lan, Jiabao Wen |
AAAI | 3 |
| 2026 | Learning-aided equivariant filtering on the special euclidean group for underwater navigation sensor fusion
Jiabao Wen, Dijing Wang, Jingyi He 0001, Meng Xi 0001, Shuai Xiao 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Nuclear Graph-Guided Multiple Instance Learning for Weakly Supervised Tumor Region DiscoveryabstractWeakly supervised tumor localization in whole-slide images (WSIs) remains challenging due to the absence of region level annotations and the difficulty of capturing diagnostically relevant cellular organization under gigapixel resolution. Most existing multiple instance learning (MIL) methods rely primarily on patch-level appearance features, overlooking the nuclear structural patterns that underlie pathological interpretation. We propose Nuclear Graph-Guided Multiple Instance Learning (NGG-MIL), a weakly supervised framework that explicitly incorporates nuclear-level structural information into WSI analysis. Within each patch, nuclei are represented as a Direction Weighted Nuclear Graph, where edges encode both spatial proximity and nuclear orientation consistency. Graph Neural Network (GNN) is employed to learn structure-aware patch representations, which are subsequently aggregated through Attention Based MIL for slide-level prediction. The learned patch attention scores are spatially reprojected to generate interpretable tumor probability maps without requiring region-level supervision. Extensive experiments on three public breast cancer WSI datasets (CAMELYON16, TCGA-BRCA, and BRACS) demonstrate that NGG-MIL consistently surpasses strong MIL baselines in both slide-level classification performance and interpretability of localization maps, achieving consistent AUC improvements and competitive F1 scores across datasets. Andi Duan, Guipeng Lan, Xiaohua Yin, Shuai Xiao 0001, Qinggang Meng, Baihua Li |
IEEE Signal Process. Lett. | 5 |
| 2025 | DWT-CPLnet: A New Intrusion Disturbance Identification Paradigm for Optical Fiber Sensing Network in Open EnvironmentsabstractPerimeter security system based on distributed optical fiber sensor network plays a key role in the monitoring and protection of restricted areas and large industrial areas. At present, most of the distributed intrusion signal recognition algorithms rely on manual feature extraction methods and traditional classifiers such as traditional support vector machines, which generally have low recognition efficiency and accuracy. To solve these problems, a convolutional prototype network DWTCPLnet is proposed in this paper. Firstly, the original onedimensional intrusion interference signal is decomposed into five approximate coefficients in the frequency domain by discrete wavelet transform (DWT), and then combined with the original signal to form a new two-dimensional data. This two-dimensional data is then entered into DWT-CPLnet for training. At the same time, the training process of the network is restricted by the metric space of prototype learning. The experimental results indicate that the average recognition accuracy of DWT-CPLnet in 6 types of common intrusion disturbance signals (three natural disturbances: wind blowing, light rain, heavy rain; three manmade disturbances: knocking, impacting and slapping) can reach 99.59%, and also has the ability to identify unknown classes to meet the actual monitoring needs. Ziqiang Huo, Meng Xi 0001, Anwer Adel Al-Dulaimi, Jiabao Wen, Shuai Xiao 0001 |
ICC | 6 |
| 2025 | Inner Information Analysis Algorithm for Deep Neural Network based on CommunityabstractDeep learning has achieved advancements across a variety of forefront fields. However, its inherent 'black box' characteristic poses challenges to the comprehension and trustworthiness of the decision-making processes within neural networks. To mitigate these challenges, we introduce InnerSightNet, an inner information analysis algorithm designed to illuminate the inner workings of deep neural networks through the perspectives of community. This approach is aimed at deciphering the intricate patterns of neurons within deep neural networks, thereby shedding light on the networks' information processing and decision-making pathways. InnerSightNet operates in three primary phases, 'neuronization-aggregation-evaluation'. Initially, it transforms learnable units into a structured network of neurons. Subsequently, these neurons are aggregated into distinct communities according to representation attributes. The final phase involves the evaluation of these communities' roles and functionalities, to unpick the information flow and decision-making. By transcending focus on single-layer or individual neuron, InnerSightNet broadens the horizon for deep neural network interpretation. InnerSightNet offers a unique vantage point, enabling insights into the collective behavior of communities within the overarching architecture, thereby enhancing transparency and trust in deep learning systems. Guipeng Lan, Shuai Xiao 0001, Meng Xi 0001, Jiabao Wen |
ICLR | 2 |
| 2025 | Enhanced deepfake detection with DenseNet and Cross-ViT
Fazeela Siddiqui, Shuai Xiao 0001, Muhammad Fahad 0013 |
Expert Syst. Appl. | 3 |
| 2025 | DD-MID: An innovative approach to assess model information discrepancy based on deep dream
Zhuo Zhang 0025, Jialin Li 0002, Shuai Xiao 0001, Jiabao Wen, Wen Lu 0005, Xinbo Gao 0001 |
Expert Syst. Appl. | 3 |
| 2025 | Diffusion model in modern detection: Advancing Deepfake techniques
Fazeela Siddiqui, Shuai Xiao 0001, Muhammad Fahad 0013 |
Knowl. Based Syst. | 3 |
| 2025 | Active Learning for Object Detection With Vectorized Dual Pseudo Loss and Multiple Instance Offset ConstraintabstractExisting active learning methods for object detection face challenges, such as the lack of ground truth labels for regression loss, insufficient representation of unlabeled instance samples information, and discrepancies in information quality between image-level and multiple anchor-level instances. To address these issues, we propose an active learning method for object detection with vectorized dual pseudo loss and multiple instance offset constraint. This method implements a two-stage framework. The first stage focuses on evaluating the information quality of detection images. We first pioneer a dual pseudo loss formulation that provides theoretically grounded regression loss estimation. The regression loss is calculated as the norm of the offset discrepancy loss vector between the enhanced and original base box vector, further constrained by the cosine value of the angle between the anchor box feature and regressor parameters vector. The distance entropy from the base box feature vector to each category's feature prototype vector is used as a weighting factor for the regression and classification information quality of instance samples. Subsequently, the second stage employs diversity-driven sampling on high-information images, leveraging instance-level cosine similarity to effectively remove redundant images. The proposed method outperforms state-of-the-art active learning approaches for object detection on PASCAL VOC and MS COCO datasets. Additionally, the proposed dual pseudo regression loss robustly captures regression information quality, demonstrating its effectiveness for active learning in object detection. Jiasai Wu, Shuai Xiao 0001, Jiabao Wen, Qinggang Meng, Wen Lu 0004, Xinbo Gao 0001 |
IEEE Trans. Cybern. | 3 |
| 2025 | Generative AI-Based Data Completeness Augmentation Algorithm for Data-Driven Smart HealthcareabstractIn the decade, artificial intelligence has achieved great popularity and applications in medicine and healthcare. Various AI-based algorithms have shown astonishing performance. However, in various data-driven smart healthcare algorithms, the problem of incomplete dataset remains a huge challenge. In this paper, we propose a data completeness enhancement algorithm based on generative AI (i.e., GenAI-DAA) to solve the problems of the in-sufficient data for model training, the data imbalance, and the biases of the training samples. We first construct the cognitive field of the generative models and effectively understand the state of incomplete cognition in generative models. Secondly, on this basis, we propose a quest algorithm for abnormal samples in the cognitive field based on local outlier factor. By fine-grained value evaluation, abnormal samples are given more refined attention. Finally, integrating the above process through multiple cognitive adjustments, GenAI-DAA gradually improves the cognitive ability. GenAI-DAA can be summarized as "Quest $ \longrightarrow$ Estimate$ \longrightarrow$Tune-up". We have conducted extensive experiments to demonstrate the effectiveness of our proposed algorithm, and shown widely applications to some typical data-driven smart healthcare algorithms. Guipeng Lan, Shuai Xiao 0001, Jiabao Wen, Meng Xi 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | MARL-Based AUV Formation for Underwater Intelligent Autonomous Transport Systems Supported by 6G NetworkabstractWith the advancement of communication technology from 5G to 6G, future communication networks will no longer be limited to land and air, and the ocean will also become the battlefield for 6G networks. The expansion of the network has expanded the scope of Intelligent Autonomous Transport Systems (IATS). As a new type of underwater transport system, Autonomous Underwater Vehicle (AUV) has gained popularity due to their advantages of autonomy, endurance, and concealment. In practical applications, it is necessary to fully consider the impact of uncertain marine environments on AUV’s motion, and also design stable control unit to achieve AUV formation. The core of the control unit is the AUV formation control algorithm, which should enable AUV to complete path planning and obstacle avoidance while ensuring formation control. In order to solve the above problems, an Intelligent Multi-agent path planning and formation control algorithm based on Value-decomposition networks (IMV) is proposed in this paper. Specifically, a three-dimensional high-resolution marine simulation environment located in the Mariana Trench is established, the state transition function and reward function are well designed under uncertain conditions for stable Multi-Agent Reinforcement Learning (MARL) mechanism, a Value-Decomposition Networks (VDN) based training framework is constructed to improve the convergence speed of the proposed method. The experimental results verify the excellent performance of the IMV method proposed in this paper, demonstrating that our method can outperform other methods in the aspect of stability, adaptability, intelligence, and timeliness. Jingyi He 0001, Meng Xi 0001, Jiabao Wen, Shuai Xiao 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | An Expert Experience-Enhanced Security Control Approach for AUVs of the Underwater Transportation Cyber-Physical SystemsabstractBy combining transportation information with physical elements, transportation cyber-physical systems (T-CPS) take advantage of the strengths of information technology and show great potential in terms of efficiency, safety, and control. T-CPS covers land, air, and underwater domains involving vehicles, drones, and autonomous underwater vehicles (AUVs), facilitating our lives and creating productivity. However, underwater T-CPS faces greater difficulties and challenges than the first two areas. On the one side, underwater equipment is generally expensive and thus requires a high level of safety. On the other side, the complexity of the marine environment causes uncertainty in the control. To address these challenges, this paper proposes an expert experience-enhanced control approach designed to enhance AUV reliability and safety. Firstly, we model AUV cluster control, including the complex underwater environment and cooperative control strategy, and refine this problem into a Markov decision problem (MDP) model based on the leader-follower strategy. Subsequently, a multi-agent reinforcement learning cluster control algorithm is developed on the framework of Centralized Training Distributed Execution (CTDE) to improve the learning and exploration capabilities of AUVs. Finally, we propose an expert experience-enhanced strategy that reduces the impact of non-smooth environments and also ameliorates the limitation of relying exclusively on rule-based experience. Experiments compare the linear and triangular AUV formation control tasks, and the proposed approach shows promising superiority and possesses sound stability in dynamically changing environments. Meng Xi 0001, Jiabao Wen, Jingyi He 0001, Shuai Xiao 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Generative Model Perception Rectification Algorithm for Trade-Off between Diversity and QualityabstractHow to balance the diversity and quality of results from generative models through perception rectification poses a significant challenge. Abnormal perception in generative models is typically caused by two factors: inadequate model structure and imbalanced data distribution. In response to this issue, we propose the dynamic model perception rectification algorithm (DMPRA) for generalized generative models. The core idea is to gain a comprehensive perception of the data in the generative model by appropriately highlighting the low-density samples in the perception space, also known as the minor group samples. The entire process can be summarized as "search-evaluation-adjustment". To identify low-density regions in the data manifold within the perception space of generative models, we introduce a filtering method based on extended neighborhood sampling. Based on the informational value of samples from low-density regions, our proposed mechanism generates informative weights to assess the significance of these samples in correcting the models' perception. By using dynamic adjustment, DMPRA ensures simultaneous enhancement of diversity and quality in the presence of imbalanced data distribution. Experimental results indicate that the algorithm has effectively improved Generative Adversarial Nets (GANs), Normalizing Flows (Flows), Variational Auto-Encoders (VAEs), and Diffusion Models (Diffusion). Guipeng Lan, Shuai Xiao 0001, Jiabao Wen |
AAAI | 2 |
| 2024 | An Konwledge-Based Semi-supervised Active Learning Method for Precision Pest Disease Diagnostic
Yong Zhu 0007, Shuai Xiao 0001, Zhuo Zhang 0025, Jiabao Wen, Meng Xi 0001 |
KSEM (1) | 2 |
| 2024 | Curvature index of image samples used to evaluate the interpretability informativeness
Zhuo Zhang 0025, Shuai Xiao 0001, Meng Xi 0001, Jiabao Wen |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Active learning inspired method in generative models
Guipeng Lan, Shuai Xiao 0001, Jiabao Wen, Wen Lu 0005, Xinbo Gao 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Face swapping with adaptive exploration-fusion mechanism and dual en-decoding tactic
Guipeng Lan, Shuai Xiao 0001, Jiabao Wen, Wen Lu 0005, Xinbo Gao 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Idea and Application to Explain Active Learning for Low-Carbon Sustainable AIoTabstractThe Internet of Things (AIoT) is supporting the revolution of many industries. However, AIoT systems require a large amount of computing resources and electricity consumption as support, which leads to significant carbon emissions and energy consumption, which is not conducive to sustainable energy development. Reducing the demand for data in artificial intelligence through active learning (AL) is an effective solution. In this study, based on the interpretability of neural networks, we propose an interpretable AL algorithm. By improving the traditional heatmap display method, we use predicted probability entropy and posterior probability entropy to form an information class activation map information visualization method, thereby providing an explanation for the sources of information in AL. Meanwhile, we propose a Similarity-Loss AL sampling strategy to evaluate the information content of samples. The experimental results show that our proposed method has achieved good results in terms of interpretability and optimization of sampling in AL. In addition, the proposed Similarity-Loss sampling strategy has achieved the highest performance in current AL scenarios, contributing to achieving low-carbon and sustainable AIoT. Shuai Xiao 0001, Meng Xi 0001, Guipeng Lan, Zhuo Zhang 0025 |
IEEE Internet Things J. | 2 |
| 2024 | A Lightweight Reinforcement-Learning-Based Real-Time Path-Planning Method for Unmanned Aerial VehiclesabstractThe Unmanned Aerial Vehicles (UAVs) are competent to perform a variety of applications, possessing great potential and promise. The Deep Neural Network (DNN) technology has enabled the UAV-assisted paradigm, accelerated the construction of smart cities, and propelled the development of the Internet of Things (IoT). UAVs play an increasingly important role in various applications, such as surveillance, environmental monitoring, emergency rescue, supplies delivery, for which a robust path planning technique is the foundation and prerequisite. However, existing methods lack comprehensive consideration of the complicated urban environment and do not provide an overall assessment of the robustness and generalization. Meanwhile, due to the resource constraints and hardware limitations of UAVs, the complexity of deploying the network needs to be reduced. This paper proposes a lightweight, reinforcement learning-based real-time path planning method for UAVs, Adaptive Soft Actor-Critic algorithm (ASAC), which optimizing training process, network architecture, and algorithmic models. First of all, we establish a framework of global training and local adaptation, where the structured environment model is constructed for interaction, and local dynamically varying information aids in improving generalization. Secondly, ASAC introduces a cross-layer connection approach that passes the original state information into the higher layers to avoid feature loss and improve learning efficiency. Finally, we propose an adaptive temperature coefficient, which flexibly adjusts the exploration probability of UAVs with the training phase and experience data accumulation. In addition, a series of comparison experiments have been conducted in conjunction with practical application requirements, and the results have fully proved the favorable superiority of ASAC. Meng Xi 0001, Huiao Dai, Jingyi He 0001, Jiabao Wen, Shuai Xiao 0001 |
IEEE Internet Things J. | 6 |
| 2024 | Emo-AEN: A Lightweight Network for Brand Image Design Based on Aesthetic Evaluation
Honglei Cheng, Haorui Yi, Guipeng Lan, Shuai Xiao 0001 |
Mob. Networks Appl. | 4 |
| 2024 | Securing the Socio-Cyber World: Multiorder Attribute Node Association Classification for Manipulated MediaabstractWith the rapid development of information technology, social network has become an indispensable part of daily life. People have been able to get news from all over the world through social networks for a long time. People spend more time online than they do in real life. However, the information we get in the world of social network is not purely benign. Due to the development of artificial intelligence technology, more and more tampered media information appears in social networks, some for entertainment, while others become the dark side of social networks, of which the most harmful is to people in the media tamper. For fake news and misinformation caused by media tampering, we need to trace the source and clearly distinguish the truth from the manipulated. This article proposes an image media forgery classification method of multiorder attribute nodes. First, we use different methods to extract the edge, texture, grayscale, and color attributes of the image. Second, according to the characteristics of different attributes, we calculate the first-order entropy of edge attributes, the second-order entropy of texture attributes, local entropy of grayscale, and color properties. Finally, we represent each image with some nodes and build a graph convolutional network (GCN) to classify real and fake images. Experimental results on mainstream media manipulation datasets show that our method is the state-of-the-art compared with similar methods. Shuai Xiao 0001, Guipeng Lan, Yang Li 0111, Jiabao Wen |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | Image Aesthetics Assessment Based on Hypernetwork of Emotion FusionabstractResearch in psychology demonstrates that visual features and semantic content can convey various emotions. Furthermore, studies have proved that image emotion and aesthetics are inextricably linked. During the image aesthetic assessment process (IAA), images elicit emotional responses from individuals, leading to emotional resonance and influencing the evaluation of images. This article proposes an image aesthetics assessment method based on hypernetwork of emotion fusion (HNEF). Our method incorporates the emotions depicted in images into the process of IAA. To accomplish this, we extract both aesthetic and emotional features from the images. Additionally, we employed the self-attention mechanism of the transformer to comprehensively investigate the intimate connection between aesthetics and emotion. Additionally, the hypernetwork is designed to establish perception rules governing the high-level semantic information in images. The experimental results validate the strong correlation between emotion and aesthetics. Furthermore, the proposed method exhibits a significantly competitive advantage when compared to existing methods on the Aesthetic Visual Analysis (AVA) dataset. Guipeng Lan, Shuai Xiao 0001, Yanshuang Zhou, Jiabao Wen, Wen Lu 0004, Xinbo Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | MCS-GAN: A Different Understanding for Generalization of Deep Forgery DetectionabstractAfter several years of development, deep synthesis technology has made significant progress in image and video synthesis. Deep forgery represented by Deepfakes has become a research hotspot, which is used as a tool for disinformation attacks. The current strongly discriminative models can have good performance on specific datasets, even close to 100% accuracy. Unfortunately, since a specific discriminative method only fits a specific data distribution, and different forgery methods or datasets have different data distributions. These methods fail to achieve high performance in cross-dataset detection. In response to this problem and focusing on the actual situation, we adjust the strong generalization detection across the dataset to the generalization detection of unseen fake video. We propose Multi-Crise-Cross Attention and StyleGANv2 Generative Adversarial Network (MCS-GAN). Firstly, we built a Generative Adversarial Network (GAN) framework to learn the distribution of real face data and generate corresponding face images. Secondly, to break the high stitch between the fake region and the background, the model needs to have strong enough feature analysis and pixel restoration capabilities. Therefore, we propose a generator consisting of a Multi-Crise-Cross-Attention (MC) encoder and a StyleGANv2 (SG2) decoder. Finally, to avoid the situation where as long as a face is normal or different faces are abnormal, we set a latent space encoding discriminator and increase the ratio of latent space vector, so as to detect anomaly generated by the forgery operation acting on latent space. We conduct some model generalization experiments on videos on the Internet and some popular deepfake databases. The results show that the accuracy of our method is better compared with the best methods. Shuai Xiao 0001, Guipeng Lan, Qinggang Meng, Xinbo Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | High Fidelity Face-Swapping With Style ConvTransformer and Latent Space SelectionabstractFace-swapping technology has been widely used in people's life, and people also put forward higher requirements for it. Most of the current face-swapping methods are difficult to generate a high-definition face image. Through StyleGAN, we can generate high-definition face images. However, face-swapping with StyleGAN is still challenging. Firstly, we need to map the target image to the latent space of StyleGAN. Many tasks need to map the input image to a new latent space for face-swapping, because identity features are complex and challenging to map to specific latent space layers directly. So face-swapping is completed in the remapping process, which consumes excess computing resources for reconstruction. And the generated image is difficult to maintain the original image color, face attributes, background and other attributes. We propose a new method, which only edits the code of w+ latent space of StyleGAN to complete the face-swapping and generate high-definition face images. We propose the GAN inversion method to improve the effect of face swapping, which combines convolution networks' advantages in extracting texture features and the benefits of transformers in extracting structure features. In the latent space of StyleGAN, the low-level feature layer is dominated by structure information, and the high-level feature layer is overwhelmed by texture information. Furthermore, we propose latent space selection, through which the neural network can learn disentangled representations of identity information in the latent space. Finally, we improved the post-processing process of face swapping to keep the image's background. Our method can complete face-swapping by editing the w+ space. Thus, high-quality face image can be generated and a lot of computing resource is saved on image reconstruction. At the same time, our method can keep other attributes better in the face-swapping process. Shuai Xiao 0001, Guipeng Lan, Jiabao Wen |
IEEE Trans. Multim. | 3 |
| 2024 | Say No to Redundant Information: Unsupervised Redundant Feature Elimination for Active LearningabstractThe usual active learning is to sample unlabeled set by designing efficient sample information evaluation algorithms. However, information redundancy between candidate sets is often overlooked. This can cause similar data to be labeled repeatedly, producing ineffective gains for the model. In this paper, we proposed an Unsupervised Redundant Feature Elimination Active Learning module (URFEAL), which utilizes the information feature coincidence of the unlabeled set to eliminate information redundant data, thus guaranteeing the validity of each candidate data. URFEAL consists of feature clusterer and eliminator. The feature clusterer computes class boundaries based on feature densities to discretize each class of the candidate set, and the eliminator judges data similarity by overlapping degree to eliminate redundant data features. Furthermore, we propose an anti-noise sampling strategy Outlier Feature Elimination (OFE) in URFEAL to filter mislabeled sets for relabeling in the data sampling stage. We extensively evaluate our method by image classification and perform experimental validation on CIFAR-10, CIFAR-100 and CALTECH-101. The experimental results show that the improvements we make are especially significant for most existing active learning algorithms in the low data stage, which demonstrates the effectiveness and generality of URFEAL. Shukun Ma, Zhuo Zhang 0025, Yang Li 0111, Shuai Xiao 0001, Jiabao Wen, Wen Lu 0005, Xinbo Gao 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Forgery Detection by Weighted Complementarity between Significant Invariance and Detail EnhancementabstractGenerative adversarial networks have shown impressive results in the modeling of movies and games, but what if such powerful image generation capability is used to harm the Multimedia? The face replacement methods represented by Deepfakes are becoming a threat to everyone, so the development of image authenticity detection methods has become a top priority. For achieving accurate detection resistant to compression effects, we propose a weighted complementary dual-stream detection method. First, to alleviate the influence of image compression on manipulation detection, we propose the concept of pixel-wise saliency invariance. We map fake images onto saliency maps via Quaternary Fourier Transform, which discovers the invariant properties of image phase spectra on different compressions. Meanwhile, to capture boundary traces more easily, we propose the concept of pixel-wise detail enhancement. We apply Bilateral Filtering to preserve the texture edges of fake images and amplify the fake boundaries. Finally, to take full advantage of the two proposed concepts, a weighted complementary dual-stream network is designed as a classifier to fuse features and identify real and fake. On different benchmarks like FaceForensics++ (FF++), Celeb-DF, and DFDC, the experimental results show that the proposed method has the average best detection accuracy compared to existing methods. Shuai Xiao 0001, Zhuo Zhang 0025, Jiabao Wen, Yang Li 0111 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Efficient data-driven behavior identification based on vision transformers for human activity understanding
Zhuo Zhang 0025, Shuai Xiao 0001, Shukun Ma, Yang Li 0111, Wen Lu 0005, Xinbo Gao 0001 |
Neurocomputing | 3 |
| 2023 | Manipulation detection of key populations under information measurement
Shuai Xiao 0001, Zhuo Zhang 0025, Jiabao Wen, Yang Li 0111 |
Inf. Sci. | 1 |
| 2023 | No-reference quality index of tone-mapped images based on authenticity, preservation, and scene expressiveness
Yang Zhao 0027, Shuai Xiao 0001, Wen Lu 0004, Xinbo Gao 0001 |
Signal Process. | 2 |
| 2022 | A controllable face forgery framework to enrich face-privacy-protection datasets
Yong Zhu 0007, Shuai Xiao 0001, Guipeng Lan, Yang Li 0111 |
Image Vis. Comput. | 3 |
| 2022 | MSTA-Net: Forgery Detection by Generating Manipulation Trace Based on Multi-Scale Self-Texture AttentionabstractLots of Deepfake videos are circulating on the Internet, which not only damages the personal rights of the forged individual, but also pollutes the web environment. What’s worse, it may trigger public opinion and endanger national security. Therefore, it is urgent to fight deep forgery. Most of the current forgery detection algorithms are based on convolutional neural networks to learn the feature differences between forged and real frames from big data. In this paper, from the perspective of image generation, we simulate the forgery process based on image generation and explore possible trace of forgery. We propose a multi-scale self-texture attention Generative Network(MSTA-Net) to track the potential texture trace in image generation process and eliminate the interference of deep forgery post-processing. Firstly, a generator with encoder-decoder is to disassemble images and performed trace generation, then we merge the generated trace image and the original map, which is input into the classifier with Resnet as the backbone. Secondly, the self-texture attention mechanism(STA) is proposed as the skip connection between the encoder and the decoder, which significantly enhances the texture characteristics in the image disassembly process and assists the generation of texture trace. Finally, we propose a loss function called Prob-tuple loss restricted by classification probability to amend the generation of forgery trace directly. To verify the performance of the MSTA-Net, we design different experiments to verify the feasibility and advancement of the method. Experimental results show that the proposed method performs well on deep forged databases represented by FaceForensics++, Celeb-DF, Deeperforensics and DFDC, and some results are reaching the state-of-the-art. Shuai Xiao 0001, Aiyun Li, Wen Lu 0005, Xinbo Gao 0001, Yang Li 0111 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Detecting fake images by identifying potential texture difference
Shuai Xiao 0001, Aiyun Li, Guipeng Lan, Huihui Wang 0001 |
Future Gener. Comput. Syst. | 2 |
| 2021 | MTD-Net: Learning to Detect Deepfakes Images by Multi-Scale Texture DifferenceabstractWith the rapid development of face manipulation technology, it is difficult for human eyes to distinguish fake face images. On the contrary, Convolutional Neural Network (CNN) discriminators can quickly reach high accuracy in identifying fake/real face images. In this study, we explore the behavior of CNN models in distinguish fake/real faces. We find multi-scale texture difference information plays an important role in face forgery detection. Motivated by the above observation, we propose a new Multi-scale Texture Difference model coined as MTD-Net for robust face forgery detection, which leverages central difference convolution (CDC) and atrous spatial pyramid pooling (ASPP). CDC combines the pixel intensity information and the pixel gradient information to give a stationary description of texture difference information. Simultaneously, based on the ASPP, multi-scale information fusion can keep the texture features from being destroyed. Experimental results on several databases, Faceforensics++, DeeperForensics-1.0, Celeb-DF and DFDC prove that our MTD-Net outperforms existing approaches. The MTD-Net is more robust to image distortion, e.g., JPEG compression and blur, which is urgently needed in the wild world. Aiyun Li, Shuai Xiao 0001, Wen Lu 0004, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |