VLDB 2026 Research / reviewers in the wild / expert
Shuaibo Li
dblp:301/9677
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-8542-4168ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SynerDetect: Hierarchical Synergistic Learning for Generalizable AI-Generated Image DetectionabstractThe rapid advancement of generative models, which produce increasingly realistic synthetic images, urgently demands robust and generalizable detection methods. Consequently, research has largely pivoted to leveraging large-scale Vision Foundation Models (VFMs) for enhanced generalization. However, existing VFM-based approaches primarily adhere to either perceptual or generative paradigms, each with limitations: perceptual models capture high-level semantics but often miss subtle artifacts, whereas generative models emphasize fine-grained flaws yet overlook semantic inconsistency. To resolve this inherent trade-off, we introduce SynerDetect, a novel hierarchical synergistic framework that fundamentally unifies the two paradigms. SynerDetect achieves deep integration of heterogeneous forensic representations through two levels of synergy: Cross-Model Interactive Distillation (CMID) distills generative forensic signals into perceptual encoders via prompt-guided reconstruction; and Optimal Transport-Guided Discriminative Contrastive Learning (OT-DCL) structurally aligns and integrates these heterogeneous representations, consolidating them into a robust, unified detection space. SynerDetect achieves superior performance on standard benchmarks (AIGCDetectBenchmark and GenImage) and attains a notable 5.20% accuracy gain on the challenging Chameleon benchmark, whose synthetic images consistently pass the Visual Turing Test. These results unequivocally validate the robust, real-world generalization of our unified cross-paradigm framework. Shuaibo Li, Zhaohu Xing, Hongqiu Wang, Pengfei Hao, Zekai Liu, Qing Zhang 0006, Lei Zhu 0003 |
AAAI | 1 |
| 2026 | RealNet: Efficient and Unsupervised Detection of AI-Generated Images via Real-Only Representation LearningabstractDetecting AI-generated images remains a persistent challenge, as existing detectors often struggle to generalize to forgeries produced by previously unseen generative models. This generalization gap mainly stems from entanglement with semantic content and overfitting to model-specific artifacts. Moreover, many state-of-the-art methods rely on large pre-trained backbones or computationally intensive pipelines, which limit their applicability in real-world, resource-constrained environments. We propose RealNet, a lightweight and unsupervised framework that constructs a disentangled, forgery-aware representation space using only real images. RealNet first extracts semantic-agnostic representations through a dual adversarial denoising mechanism, producing compact features with low intra-class variance. These representations are then perturbed in feature space to generate pseudo-negative samples, which are combined with the original real features to train a lightweight discriminator, enabling robust detection without any dependence on synthetic images during training. Comprehensive evaluations across GAN, diffusion, and emerging VAR-based paradigms demonstrate that RealNet achieves superior cross-model generalization and robustness. RealNet surpasses previous state-of-the-art approaches by 4.51% in accuracy and 3.93% in average precision, while maintaining significantly lower computational cost. Furthermore, we introduce a medically relevant synthetic image dataset and show RealNet remains effective under severe distribution shifts, highlighting its potential for deployment in high-stakes real-world scenarios. Together, these advantages position RealNet as a practical, scalable and socially impactful solution for robust AI-generated image detection. Shuaibo Li, Laixin Zhang, Wei Ma 0008, Jianwei Guo 0003, Shibiao Xu, Zhijie Qiu, Hongbin Zha |
AAAI | 1 |
| 2026 | DualScope: Capturing Critical Spatial and Temporal Cues for Distracted Driving Activity RecognitionabstractAccurately recognizing distracted driving activities in real-world scenarios is essential for improving road and pedestrian safety. However, existing approaches are prone to attending to irrelevant scene context and are susceptible to interference from redundant frames, compromising their robustness in complex driving environments. To overcome these limitations, we propose DualScope, a novel framework that captures behaviorally critical information from both spatial and temporal perspectives. In the spatial domain, we introduce a Synergistic Behavior-Centric Distillation mechanism that leverages two key information sources: (1) position-aware knowledge derived from the SAM model, which enhances the perception of critical regions and their semantic interaction structures; and (2) fine-grained visual details obtained from cropped key regions, which improve the model's ability to capture detailed patterns within behavior-relevant areas. In the temporal domain, we present the Saliency-Aware Fine-to-Coarse Temporal Modeling module, comprising three components: a Fine-Grained Motion Encoder for capturing local inter-frame dependencies; a Dynamic Difference Extractor for generating salient motion dynamics; and a Saliency-Aware Temporal Pyramid Mamba for integrating these representations to enable multi-scale temporal modeling. This design effectively captures both short-term motions and long-term behavioral patterns. Furthermore, incorporating salient dynamics enhances the model's focus on significant behavioral variations. Extensive experiments on seven publicly available DDAR datasets demonstrate that DualScope consistently outperforms state-of-the-art methods, validating its effectiveness in capturing behavioral cues across spatial and temporal dimensions. Zhijie Qiu, Shuaibo Li, Laixin Zhang, Xuming Hu, Wei Ma 0008 |
AAAI | 2 |
| 2025 | Surgical-MambaLLM: Mamba2-Enhanced Multimodal Large Language Model for VQLA in Robotic Surgery
Pengfei Hao, Hongqiu Wang, Shuaibo Li, Zhaohu Xing, Guang Yang 0006, Kaishun Wu, Lei Zhu 0003 |
MICCAI (9) | 3 |
| 2025 | Toward Medical Deepfake Detection: A Comprehensive Dataset and Novel Method
Shuaibo Li, Zhaohu Xing, Hongqiu Wang, Pengfei Hao, Zekai Liu, Lei Zhu 0003 |
MICCAI (14) | 1 |
| 2025 | Multi-target opinion words extraction
Zixue Zhao, Shuaibo Li, Zhengpeng Li, Kejin Li |
Appl. Intell. | 2 |
| 2024 | UnionFormer: Unified-Learning Transformer with Multi-View Representation for Image Manipulation Detection and LocalizationabstractWe present UnionFormer, a novel framework that inte-grates tampering clues across three views by unified learning for image manipulation detection and localization. Specifically, we construct a BSFI-Net to extract tampering features from RGB and noise views, achieving enhanced responsive-ness to boundary artifacts while modulating spatial consis-tency at different scales. Additionally, to explore the incon-sistency between objects as a new view of clues, we combine object consistency modeling with tampering detection and localization into a three-task unified learning process, allowing them to promote and improve mutually. Therefore, we acquire a unified manipulation discriminative representation under multi-scale supervision that consolidates information from three views. This integration facilitates highly effective concurrent detection and localization of tampering. We perform extensive experiments on diverse datasets, and the results show that the proposed approach outperforms state-of-the-art methods in tampering detection and localization. Shuaibo Li, Wei Ma 0008, Jianwei Guo 0003, Shibiao Xu, Benchong Li, Xiaopeng Zhang 0001 |
CVPR | 1 |
| 2023 | A Topic Inference Chinese News Headline Generation Method Integrating Copy Mechanism
Zhengpeng Li, Jiawei Miao, Xinmiao Yu, Shuaibo Li |
Neural Process. Lett. | 5 |
| 2023 | Image Manipulation Localization Using Attentional Cross-Domain CNN FeaturesabstractAlong with the advancement of manipulation technologies, image modification is becoming increasingly convenient and imperceptible. To tackle the challenging image tampering detection problem, this article presents an attentional cross-domain deep architecture, which can be trained end-to-end. This architecture is composed of three convolutional neural network (CNN) streams to extract three types of features, including visual perception, resampling, and local inconsistency features, from spatial and frequency domains. The multitype and cross-domain features are then combined to formulate hybrid features to distinguish manipulated regions from nonmanipulated parts. Compared with other deep architectures, the proposed one spans a more complementary and discriminative feature space by integrating richer types of features from different domains in a unified end-to-end trainable framework and thus can better capture artifacts caused by different types of manipulations. In addition, we design and train a module called tampering discriminative attention network (TDA-Net) to highlight suspicious parts. These part-level representations are then integrated with the global ones to further enhance the discriminating capability of the hybrid features. To adequately train the proposed architecture, we synthesize a large dataset containing various types of manipulations based on DRESDEN and COCO. Experiments on four public datasets demonstrate that the proposed model can localize various manipulations and achieve the state-of-the-art performance. We also conduct ablation studies to verify the effectiveness of each stream and the TDA-Net module. Shuaibo Li, Shibiao Xu, Wei Ma 0008, Qiu Zong |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | MHDP: An Efficient Data Lake Platform for Medical Multi-source Heterogeneous Data
Peng Ren 0005, Shuaibo Li, Wenkui Zheng, Qin Cui, Wang Chang, Xin Li 0111, Chun Zeng, Ming Sheng, Yong Zhang 0002 |
WISA | 2 |
| 2021 | Intelligent Visualization System for Big Multi-source Medical Data Based on Data Lake
Peng Ren 0005, Ziyun Mao, Shuaibo Li, Yating Ke, Lanyu Yao, Xin Li 0111, Ming Sheng, Yong Zhang 0002 |
WISA | 3 |