VLDB 2026 Research / reviewers in the wild / expert
Jiabao Guo
dblp:202/3909
· DBLP profile ↗
20ranked-venue papers
2as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PAGPL: Privacy-Aware Graph Prompt Learning Scheme via Adaptive Perturbation-Estimated Topology RecoveryabstractGraph prompt learning (GPL) serves as a crucial framework for mitigating the knowledge transfer by reconciling the substantial mismatch between pre-training models and downstream tasks. However, prevalent GPL paradigm fail to accommodate graph data affected by privacy-induced noise. Specifically, 1) GPL typically relies on the stability of original graph structures for the design of effective prompt templates; 2) the construction of prompts lacks explicit guidance to suppress noise introduced by privacy perturbations; 3) prompt optimization on single disturbed graphs can easily lead to overfitting to noise patterns. To address these issues, we propose a novel privacy-aware graph prompt learning (PAGPL) scheme, which alleviates spurious clues caused by privacy noise injection. Initially, an adaptive structure-wise Bayesian estimation is applied to reconstruct the privacy-perturbed graphs. Subsequently, to suppress the impact of residual perturbation, a noise-resilient prompt generation is employed to filter unreliable structural and signals. Ultimately, we incorporate a multi-view-based progressive privacy consistency to promote the robustness of prompts against the semantic misalignment while improving the task-specific consistency. The experimental results reveal that our scheme outperforms state-of-the-art (SOTA) GPL approaches with a 10%–60% improvement in accuracy under various real-world privacy-perturbed scenarios. Ju Jia, Jiansen Song, Jingxuan Yu, Jiabao Guo, Xiaoshuang Jia, Di Wu 0050, Yali Yuan, Guang Cheng 0001 |
AAAI | 4 |
| 2026 | Wi-CBR: Salient-aware Adaptive WiFi Sensing for Cross-domain Behavior RecognitionabstractThe challenge in WiFi-based cross-domain Behavior Recognition lies in the significant interference of domain-specific signals on gesture variation. However, previous methods alleviate this interference by mapping the phase from multiple domains into a common feature space. If the Doppler Frequency Shift (DFS) signal is used to dynamically supplement the phase features to achieve better generalization, it enables the model to not only explore a wider feature space but also to avoid potential degradation of gesture semantic information. Specifically, we propose a novel Salient-aware Adaptive WiFi Sensing for Cross-domain Behavior Recognition (Wi-CBR), which constructs a dual-branch self-attention module that captures temporal features from phase information reflecting dynamic path length variations while extracting kinematic features from DFS correlated with motion velocity. Moreover, we design a Saliency Guidance Module that employs group attention mechanisms to mine critical activity features and utilizes gating mechanisms to optimize information entropy, facilitating feature fusion and enabling effective interaction between salient and non-salient behavioral characteristics. Extensive experiments on two large-scale public datasets (Widar3.0 and XRF55) demonstrate the superior performance of our method in both in-domain and cross-domain scenarios. Ruobei Zhang, Shengeng Tang, Xiang Zhang 0011, Jiabao Guo |
AAAI | 5 |
| 2026 | CTEA: Camouflaged topological element attack via causal influence discovery
Ju Jia, Pengyuan Gao, Meng Luo 0002, Cong Wu 0003, Jiabao Guo |
Expert Syst. Appl. | 5 |
| 2026 | ICPE-FAS: Instance and Category Prompts Engineering for Generalizable Face Anti-Spoofing
Ajian Liu 0001, Xun Lin, Hui Ma 0018, Xinxing Yu, Jiabao Guo, Zitong Yu, Jun Wan 0001, Zhanchuan Cai, Zhen Lei 0001, Yanyan Liang 0001 |
Int. J. Comput. Vis. | 5 |
| 2026 | Take off Your Disguise: Detecting Disguised Prompt-Based Jailbreak Attacks Against LLMsabstractLarge language models (LLMs) are typically equipped with alignment mechanisms designed to prevent the generation of harmful content. However, recent advances have led to the emergence of highly evasive disguised prompt jailbreak attacks (DPJAs), where attackers conceal real malicious intent within prompts that appear benign on the surface, thereby inducing the model to produce unsafe outputs. In this work, we propose prompt reconstruction-based detection (RePrompt), a training-free and plug-in detection framework designed to identify such jailbreak attacks. The key insight of RePrompt is that, when guided by carefully designed templates, LLMs possess the capability to uncover disguised adversarial prompts and effectively detect jailbreak attacks. RePrompt leverages the reasoning capabilities of the LLM itself to analyze and uncover the true intent behind a user’s query. When a disguised jailbreak prompt is encountered, RePrompt reconstructs the underlying malicious intent hidden beneath the surface-level disguise. Experiments across multiple LLMs, datasets, and three representative DPJAs demonstrate that RePrompt consistently outperforms state-of-the-art defense methods, reducing the average attack success rate from 47.0% to below 1.6%, achieving a near-zero false positive rate, and limiting the average query cost to only 1.13. Hui Liu 0018, Fujv Wen, Hongqin Du, Jiabao Guo, Bo Zhao 0023 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2026 | TG4MM: Time-Varying Gaussian Splatting for 3D Motion Magnificationabstract3D motion magnification aims to enable us to visualize subtle, imperceptible motions by integrating eulerian video magnification with novel view synthesis. Existing method extracts the variation of feature embeddings using Neural Radiance Fields (NeRF) over time. However, this volume rendering technique suffers from two shortcomings for 3D motion magnification: (1) When reconstructing time-varying scenes through volume rendering, spatial-temporal operations between static and dynamic representations often generate noticeable artifacts, leading to blurred magnified frames. (2) When processing high-resolution dynamic scenes, the intrinsically low rendering efficiency of these techniques causes excessive computational latency, preventing real-time visualization. In this work, instead of NeRF, we propose a novelTime-varying Gaussian Splatting for 3D Motion Magnification(TG4MM) that is capable of achieving real-time rendering while effectively handling blurred magnified frames in dynamic 3D motion magnification scenes. Specifically, we propose a motion-space decoupled triplane modeling approach. The space triplane captures major spatial structures from the first frame, while the motion triplane captures subtle motion information from subsequent frames. Furthermore, we develop a phase-based motion magnification module that enhances subtle motions by applying filters within the embedding space and subtle motion triplane. Experimental results demonstrate the effectiveness of our method, showing that it outperforms existing 3D motion magnification techniques and achieves a speed up to 126 FPS. Jiabao Guo, Fei Wang 0073, Jinyang Huang, Zhi Liu 0002, Dan Guo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | SMInject: Specious Malignant Injection Attacks With Semantically-Enhanced Tokens in Cross-Modal RetrievalabstractThe pre-training multimodal models have achieved remarkable success with powerful cross-modal understanding capabilities, while easily being affected by deliberate injection attacks. Although the deceptive injection attacks are harmful, they are valuable in revealing the vulnerability and improving the robustness for multimodal models. Unfortunately, the existing multimodal injection attacks pay less attention to the complicated roles of different modality-related causal correlation, which results in such attacks being susceptible to detection and defense. To alleviate this issue, we propose a novel specious malignant injection attack framework, calledSMInject, which exploits both the irrationality and causal correlation across diverse modalities to stealthily manipulate the space of output. To enhance the stealthiness, we generate deceptive injections to assemble the concepts by analyzing causal correlation under four types of attacks. To further boost the effectiveness, the malignant injections are guided to penetrate in the encoded embedding space by designing the premise-hypothesis consensus alignment. Extensive experiments on representative multimodal models demonstrate that ourSMInjectachieves over 14% higher attack success rate and 6% higher Hit@5 metric than state-of-the-art methods while preserving the overall utility of models. Moreover, we highlight that theSMInjectalso exhibits the desired transferability by investigating the impact of contextual factors, such as similar attack profiles, imperceptible noise perturbations,etc. Our code is available athttps://anonymous.4open.science/r/SMInject-0DBC. Ju Jia, Jiabao Guo, Xiaojun Jia, Siqi Ma 0001, Jie Gui, Robert H. Deng |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2026 | Adversarial Face Database against Deep Learning-Enabled Reconstruction AttacksabstractFace recognition systems offer a range of applications that enhance security, efficiency, and personalization, e.g., access control, identity verification, and personalized services. Mainstream facial recognition systems employ the Edge-Cloud architecture to protect user privacy by storing facial feature data instead of original facial images. However, recently emerging reconstruction attacks based on deep learning can recover the visual information of original facial images from facial features, resulting in face privacy disclosure. Existing anti-reconstruction approaches either compromise facial recognition accuracy or fail to meet real-time requirements. In this article, we propose a practical privacy-preserving approach based on adversarial perturbations against reconstruction attacks. By incorporating subtle adversarial interference into facial features, the mapping relationship from facial features to original facial images is disrupted, and the baseline reconstruction networks cannot recover the original face image. We conducted experiments on two facial recognition models, FaceNet and ArcFace, both widely deployed in practical scenarios. The results show that the face recognition accuracy sacrifice of less than 1% can significantly reduce the quality of the reconstructed image. In terms of efficiency, the average time to generate an adversarial facial feature is less than 10 ms, meeting the real-time requirements of facial recognition. Hui Liu 0018, Jiageng Chen, Jiabao Guo |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2026 | Domain Generalization for Face Anti-Spoofing via Content-Aware Composite Prompt EngineeringabstractThe challenge of Domain Generalization (DG) in Face Anti-Spoofing (FAS) is the significant interference of domain-specific signals on subtle spoofing clues. Recently, some CLIP-based algorithms have been developed to alleviate this interference by adjusting the weights of visual classifiers. How-ever, our analysis of this class-wise prompt engineering suffers from two shortcomings for DG FAS: (1) The categories of facial categories, such as real or spoof, have no semantics for the CLIP model, making it difficult to learn accurate category descriptions. (2) A single form of prompt cannot portray the various types of spoofing. In this work, instead of class-wise prompts, we propose a novel Content-aware Composite Prompt Engineering (CCPE) that generates instance-wise composite prompts, including both fixed template and learnable prompts. Specifically, our CCPE constructs content-aware prompts from two branches: (1) Inherent content prompt explicitly benefits from abundant transferred knowledge from the instruction-based Large Language Model (LLM). (2) Learnable content prompts implicitly extract the most informative visual content via Q-Former. Moreover, we design a Cross-Modal Guidance Module (CGM) that dynamically adjusts unimodal features for fusion to achieve better generalized FAS. Finally, our CCPE has been validated for its effectiveness in multiple cross-domain experiments and achieves state-of-the-art (SOTA) results. Jiabao Guo, Ajian Liu 0001, Yunfeng Diao, Hui Ma 0018, Bo Zhao 0023, Richang Hong, Meng Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | SUEDE: Shared Unified Experts for Physical- Digital Face Attack Detection EnhancementabstractFace recognition systems are vulnerable to physical attacks (e.g., printed photos) and digital threats (e.g., DeepFake), which are currently being studied as independent visual tasks, such as Face Anti-Spoofing and Forgery Detection. The inherent differences among various attack types present significant challenges in identifying a common feature space, making it difficult to develop a unified framework for detecting data from both attack modalities simultaneously. Inspired by the efficacy of Mixture-of-Experts (MoE) in learning across diverse domains, we explore utilizing multiple experts to learn the distinct features of various attack types. However, the feature distributions of physical and digital attacks overlap and differ. This suggests that relying solely on distinct experts to learn the unique features of each attack type may overlook shared knowledge between them. To address these issues, we propose SUEDE, the Shared Unified Experts for Physical-Digital Face Attack Detection Enhancement. SUEDE combines a shared expert (always activated) to capture common features for both attack types and multiple routed experts (selectively activated) for specific attack types. Further, we integrate CLIP as the base network to ensure the shared expert benefits from prior visual knowledge and align visual-text representations in a unified space. Extensive results demonstrate SUEDE achieves superior performance compared to state-of-the-art unified detection methods. Zuying Xie, Changtao Miao, Ajian Liu 0001, Jiabao Guo, Feng Li 0037, Dan Guo 0001, Yunfeng Diao |
ICME | 4 |
| 2025 | A cutting-edge framework for industrial intrusion detection: Privacy-preserving, cost-friendly, and powered by federated learning
Lingzi Zhu, Bo Zhao 0023, Jiabao Guo, Minzhi Ji, Junru Peng |
Appl. Intell. | 3 |
| 2025 | Exploiting Non-Collinear Array Geometry for Channel Phase Error Self-CalibrationabstractThis letter proposes a novel self-calibration method for channel phase error (CPE) in antenna arrays by leveraging the geometric properties of non-collinear configurations. The CPE can be decomposed into linear and orthogonal components relative to the baseline length vector, which cause direction-of-arrival (DOA) bias and manifold distortion, respectively. However, for self-calibration methods, only the manifold distortion can be perceived and corrected. In collinear arrays, the DOA bias is independent of manifold distortion, so the linear component of CPE can't be self-calibrated. Fortunately, in non-collinear arrays, we can couple the DOA bias into the manifold distortion by reformulating the slant range formula with the Fresnel and virtual collinear array (VCA) approximations, making it possible to estimate the entire CPE without external references. We then propose an iterative least-squares algorithm that corrects for CPE using multiple snapshots collected from a non-collinear array. Simulations under various SNRs, distances, linear components of CPE, degrees of non-collinearity, and fields of view demonstrate the effectiveness and robustness of the proposed method. Qiancheng Yan, Xiaolan Qiu, Jiabao Guo, Zekun Jiao, Chibiao Ding |
IEEE Signal Process. Lett. | 3 |
| 2024 | Style-conditional Prompt Token Learning for Generalizable Face Anti-spoofingabstractFace anti-spoofing (FAS) based on domain generalization (DG) has attracted increasing attention from researchers.The reason for the poor generalization is that the model is overfitted to salient liveness-irrelevant signals.However, the previous methods alleviate the overfitting by mapping the images from multiple domains into a common feature space or promoting the separation of image features from domain-specific features and task-related features.If the text features of vision-language pre-trained (VLP) models (e.g., CLIP) are used to dynamically adjust the image features to gain a better generalization, we can not only explore a wider feature space but also avoid the potential degradation of semantic information.Specifically, we propose a FAS method of Style-Conditional Prompt Token Learning (S-CPTL), which aims to generate generalized text features by training the introduced prompt tokens to carry visual styles and use them as weights for classifiers to improve the model's generalization.Compared to the inherently static prompt token, we propose the dynamic prompt token, which can adaptively capture live-irrelevant signals from the instance-specific styles and increase their diversity through mixed feature statistics to further reduce the overfitting of the model.Thorough experimental analysis demonstrates that S-CPTL exceeds current top-performing methods in four distinct cross-dataset benchmarks. Jiabao Guo, Huan Liu 0030, Yizhi Luo, Xueli Hu, Hang Zou 0002, Yuan Zhang 0023, Hui Liu 0018, Bo Zhao 0023 |
ACM Multimedia | 1 |
| 2024 | A lightweight unsupervised adversarial detector based on autoencoder and isolation forest
Hui Liu 0018, Bo Zhao 0023, Jiabao Guo, Kehuan Zhang, Peng Liu 0005 |
Pattern Recognit. | 3 |
| 2023 | FedCSS: Joint Client-and-Sample Selection for Hard Sample-Aware Noise-Robust Federated LearningabstractFederated Learning (FL) enables a large number of data owners (a.k.a. FL clients) to jointly train a machine learning model without disclosing private local data. The importance of local data samples to the FL model vary widely. This is exacerbated by the presence of noisy data, which exhibit large losses similar to important (hard) samples. Currently, there lacks an FL approach that can effectively distinguish hard samples (which are beneficial) from noisy samples (which are harmful). To bridge this gap, we propose the Federated Client and Sample Selection (FedCSS) approach. It is a bilevel optimization approach for FL client-and-sample selection to achieve hard sample-aware noise-robust learning in a privacy preserving manner. It performs meta-learning based online approximation to iteratively update global FL models, select the most positively influential samples and deal with training data noise. Theoretical analysis shows that it is guaranteed to converge in an efficient manner. Experimental comparison against six state-of-the-art baselines on five real-world datasets in the presence of data noise and heterogeneity shows that it achieves up to 26.4% higher test accuracy, while saving communication and computation costs by at least 41.5% and 1.2%, respectively. Anran Li 0001, Jiabao Guo, Hongyi Peng, Qing Guo 0005, Han Yu 0001 |
Proc. ACM Manag. Data | 3 |
| 2023 | Cellular Traffic Prediction: A Deep Learning Method Considering Dynamic Nonlocal Spatial Correlation, Self-Attention, and Correlation of Spatiotemporal Feature FusionabstractCellular traffic prediction will play a key role in the deployment of future smart cities. Although the current traffic prediction methods based on deep learning show better performance than traditional prediction methods, they still have the following problems: (1) In spatial domain, the correlations between cellular traffic features cannot be captured accurately in non-local (including “geographic adjacency” and long-distance) spatial areas. (2) In temporal domain, the correlation of different time-grained features is failed to consider. To address these problems, a deep learning method considering dynamic non-local spatial correlation, self-attention, and correlation of spatio-temporal feature fusion is proposed. In spatial domain, our method can accurately capture the spatial correlation and highlight the contribution of more relevant traffic in the non-local area by designing a NLG-NLAM model. In temporal domain, the correlations of time-periodic features with different granularities are considered to clarify the key roles of different periodic features and eliminate the influence of irrelevant cellular traffic features on the prediction by designing a calibration layer. Experimental results indicate that the proposed method shows better performance than other mainstream prediction methods on three real-world cellular traffic datasets. Zheheng Rao, Yanyan Xu 0003, Shaoming Pan, Jiabao Guo, Yuejing Yan |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2023 | DACAS: integration of attribute-based access control for northbound interface security in SDN
Jiabao Guo |
World Wide Web (WWW) | 4 |
| 2020 | FoolChecker: A platform to evaluate the robustness of images against adversarial attacks
Hui Liu 0018, Bo Zhao 0014, Linquan Huang, Jiabao Guo |
Neurocomputing | 4 |
| 2019 | Bidirectional LSTM with attention mechanism and convolutional layer for text classification
Gang Liu 0029, Jiabao Guo |
Neurocomputing | 2 |
| 2017 | Quantitative Deliberation Model and the Method of Consensus Building
Caiquan Xiong, Jiabao Guo, Gang Liu 0029 |
CISIS | 3 |