VLDB 2026 Research / reviewers in the wild / expert
Xiao Pu 0002
dblp:91/4650-2
· DBLP profile ↗
15ranked-venue papers
6as first author
15since 2021 · last 2026
0009-0005-4173-4996ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TGDD: Trajectory Guided Dataset Distillation with Balanced DistributionabstractDataset distillation compresses large datasets into compact synthetic ones to reduce storage and computational costs. Among various approaches, distribution matching (DM)-based methods have attracted attention for their high efficiency. However, they often overlook the evolution of feature representations during training, which limits the expressiveness of synthetic data and weakens downstream performance. To address this issue, we propose Trajectory Guided Dataset Distillation (TGDD), which reformulates distribution matching as a dynamic alignment process along the model’s training trajectory. At each training stage, TGDD captures evolving semantics by aligning the feature distribution between the synthetic and original dataset. Meanwhile, it introduces a distribution constraint regularization to reduce class overlap. This design helps synthetic data preserve both semantic diversity and representativeness, improving performance in downstream tasks. Without additional optimization overhead, TGDD achieves a favorable balance between performance and efficiency. Experiments on ten datasets demonstrate that TGDD achieves state-of-the-art performance, notably a 5.0% accuracy gain on high-resolution benchmarks. Fengli Ran, Xiao Pu 0002, Bo Liu 0047, Xiuli Bi, Bin Xiao 0002 |
AAAI | 2 |
| 2026 | Breaking the Generator Barrier: Disentangled Representation for Generalizable AI-Text DetectionabstractAs large language models (LLMs) generate text that increasingly resembles human writing, the subtle cues that distinguish AI-generated content from human-written content become increasingly challenging to capture. Reliance on generator-specific artifacts is inherently unstable, since new models emerge rapidly and reduce the robustness of such shortcuts. This generalizes unseen generators as a central and challenging problem for AI-text detection. To tackle this challenge, we propose a progressively structured framework that disentangles AI-detection semantics from generator-aware artifacts. This is achieved through a compact latent encoding that encourages semantic minimality, followed by perturbation-based regularization to reduce residual entanglement, and finally a discriminative adaptation stage that aligns representations with task objectives. Experiments on MAGE benchmark, covering 20 representative LLMs across 7 categories, demonstrate consistent improvements over state-of-the-art methods, achieving up to 24.2% accuracy gain and 26.2% F_1 improvement. Notably, performance continues to improve as the diversity of training generators increases, confirming strong scalability and generalization in open-set scenarios. Our source code will be publicly available at https://github.com/PuXiao06/DRGD. Xiao Pu 0002, Zepeng Cheng, Lin Yuan 0002, Yu Wu 0001, Xiuli Bi |
ACL (1) | 1 |
| 2025 | GADNet: Improving image-text matching via graph-based aggregation and disentanglement
Xiao Pu 0002, Lin Yuan 0002, Yu Wu 0001, Liping Jing, Xinbo Gao 0001 |
Pattern Recognit. | 1 |
| 2025 | DEAR: Disentangled Event-Agnostic Representation Learning for Early Fake News DetectionabstractAbstract Detecting fake news early is challenging due to the absence of labeled articles for emerging events in training data. To address this, we propose a Disentangled Event-Agnostic Representation (DEAR) learning approach. Our method begins with a BERT-based adaptive multi-grained semantic encoder that captures hierarchical and comprehensive textual representations of the input news content. To effectively separate latent authenticity-related and event-specific knowledge within the news content, we employ a disentanglement architecture. To further enhance the decoupling effect, we introduce a cross-perturbation mechanism that perturbs authenticity-related representation with the event-specific one, and vice versa, deriving a robust and discerning authenticity-related signal. Additionally, we implement a refinement learning scheme to minimize potential interactions between two decoupled representations, ensuring that the authenticity signal remains strong and unaffected by event-specific details. Experimental results demonstrate that our approach effectively mitigates the impact of event-specific influence, outperforming state-of-the-art methods. In particular, it achieves a 6.0% improvement in accuracy on the PHEME dataset over MDDA, a similar approach that decouples latent content and style knowledge, in scenarios involving articles from unseen events different from the topics of the training set. Xiao Pu 0002, Xiuli Bi, Yu Wu 0001, Xinbo Gao 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2024 | Towards deep understanding of graph convolutional networks for relation extraction
Tao Wu 0003, Xiaolin You, Xingping Xian, Xiao Pu 0002, Shaojie Qiao, Chao Wang 0025 |
Data Knowl. Eng. | 4 |
| 2024 | Multiview-Ensemble-Learning-Based Robust Graph Convolutional Networks Against Adversarial AttacksabstractGraph neural networks (GNNs) have been widely applied in the Internet of Things (IoT) for the intelligent analysis of data collected by sensors, particularly complex relationships and dependent information between IoT devices. However, recent studies have shown that GNNs are vulnerable to adversarial attacks, which significantly limits their application in safety-critical IoT systems such as smart health monitoring, traffic monitoring, and autonomous driving. To address this issue, in addition to the low feature similarity, this study examines the vulnerability of GNNs empirically and reveals that adversarial perturbations against GNNs tend to have low structural proximity in local neighborhoods. Thus, a natural approach for defending GNNs against adversarial attacks is to utilize the related high-order robust information of the perturbed graphs. In this study, we construct auxiliary views with high-order structure and feature similarity from a perturbed graph and propose a multi-view ensemble learning-based robust graph convolutional network (MV-RGCN). Each base model in the MV-RGCN aggregates the adversarial perturbed graph and the constructed view through an adaptive aggregation mechanism, thereby eliminating the impact of adversarial perturbations. Robust representations of the base models are then integrated using an adaptive ensemble mechanism to generate predictions. Extensive experiments under adversarial attack scenarios demonstrate that the MV-RGCN outperforms state-of-the-art methods and can achieve satisfactory performance without affecting its accuracy on the original graph data. This code is available at https://github.com/thomaslok0516/MVRGCN. Tao Wu 0003, Junhui Luo, Shaojie Qiao, Chao Wang 0025, Lin Yuan 0002, Xiao Pu 0002, Xingping Xian |
IEEE Internet Things J. | 6 |
| 2024 | MiC: Image-text Matching in Circles with cross-modal generative knowledge enhancement
Xiao Pu 0002, Lin Yuan 0002, Yan Zhang 0108, Liping Jing, Xinbo Gao 0001 |
Knowl. Based Syst. | 1 |
| 2024 | Improving Image-Text Matching by Integrating Word Sense DisambiguationabstractThis letter presents a novel approach to enhance image-text matching by incorporating word sense disambiguation (WSD) within the text encoder. Our method explicitly models the senses of potentially ambiguous words, refining the semantic understanding between images and text. We introduce a sense-aware mechanism for image-text alignment by integrating a lightweight WSD component into the matching framework, optimizing both tasks simultaneously. Our WSD module operates on extensive word contexts, leveraging the power of graph attention networks (GAT), and distills knowledge from a substantially larger pre-trained WSD model through multi-task learning. Our experiments demonstrate the effectiveness of augmenting original word embeddings with sense representations derived from our WSD approach. We systematically evaluate our method against several baselines and state-of-the-art approaches on two widely-used image-text matching benchmarks: MS-COCO and Flickr30K. The results illustrate significant improvements in matching accuracy, highlighting the efficacy of our proposed approach. Xiao Pu 0002, Lin Yuan 0002, Xinbo Gao 0001 |
IEEE Signal Process. Lett. | 1 |
| 2024 | Invertible Image Obfuscation for Facial Privacy Protection via Secure FlowabstractThis paper presents a fresh paradigm for protecting facial privacy via an invertible image obfuscation framework that incorporates multiple characteristics including anonymity, diversity, reversibility, security, and lightweight all at once. We name the framework PRO-Face S, an acronym for Privacy-preserving Reversible Obfuscation of Face images via Secure flow. The core of the proposed framework is a flow-based generative model (or invertible neural network), which takes as input a face image along with its pre-obfuscated form, and outputs the privacy-protected image that visually mirrors the pre-obfuscated one. The pre-obfuscation applied can be in various forms with different types and strengths. The invertibility of the flow-based model ensures that the original image can be easily recovered from the protected image in high fidelity. An elaborate secret key mechanism is devised to securely guide the mutual transformations of privacy protection and image recovery, such that the correct recovery is only possible upon the availability of the correct secret, pre-specified by the user in the protection stage. Two modes of wrong recovery are investigated to deal with malicious recovery attempts in different scenarios. Finally, extensive experiments conducted on multiple image datasets demonstrate the superiority of the proposed framework over state-of-the-art methods. Lin Yuan 0002, Xiao Pu 0002, Yan Zhang 0108, Jiaxu Leng, Tao Wu 0003, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | PRO-Face C: Privacy-Preserving Recognition of Obfuscated Face via Feature CompensationabstractThe advancement of face recognition technology has delivered substantial societal advantages. However, it has also raised global privacy concerns due to the ubiquitous collection and potential misuse of individuals’ facial data. This presents a notable paradox: while there is a societal demand for a robust face recognition ecosystem to ensure public security and convenience, an increasing number of individuals are hesitant to release their facial data. Numerous studies have endeavored to find such a utility-privacy trade-off, yet many struggle with the dilemma of prioritizing one at the expense of the other. In response to this challenge, this paper proposes PRO-Face C, a novel paradigm for privacy-preserving recognition of obfuscated faces via a dedicated feature compensation mechanism, aimed at optimizing the equilibrium between privacy preservation and utility maximization. The proposed approach is characterized by a specialized client-server architecture: the client transmits only obfuscated images to the server, which then performs identity recognition using a pre-trained model in conjunction with a suite of privacy-free complementary features. This framework facilitates accurate face identification while safeguarding the original facial appearance from explicit disclosure. Furthermore, the obfuscated image retains its visualization capability, crucial for image preview functionalities. To ensure the desired properties, we have developed an identity-guided feature compensation mechanism, complemented by several privacy-enhancing techniques. Extensive experiments conducted across multiple face datasets underscore the effectiveness of the proposed approach in diverse scenarios. Lin Yuan 0002, Xiao Pu 0002, Yan Zhang 0108, Yushu Zhang 0001, Xinbo Gao 0001, Touradj Ebrahimi |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Contextual Learning in Fourier Complex Field for VHR Remote Sensing ImagesabstractVery high-resolution (VHR) remote sensing (RS) image classification is the fundamental task for RS image analysis and understanding. Recently, Transformer-based models demonstrated outstanding potential for learning high-order contextual relationships from natural images with general resolution ( pixels) and achieved remarkable results on general image classification tasks. However, the complexity of the naive Transformer grows quadratically with the increase in image size, which prevents Transformer-based models from VHR RS image ( pixels) classification and other computationally expensive downstream tasks. To this end, we propose to decompose the expensive self-attention (SA) into real and imaginary parts via discrete Fourier transform (DFT) and, therefore, propose an efficient complex SA (CSA) mechanism. Benefiting from the conjugated symmetric property of DFT, CSA is capable to model the high-order contextual information with less than half computations of naive SA. To overcome the gradient explosion in Fourier complex field, we replace the Softmax function with the carefully designed Logmax function to normalize the attention map of CSA and stabilize the gradient propagation. By stacking various layers of CSA blocks, we propose the Fourier complex Transformer (FCT) model to learn global contextual information from VHR aerial images following the hierarchical manners. Universal experiments conducted on commonly used RS classification datasets demonstrate the effectiveness and efficiency of FCT, especially on VHR RS images. The source code of FCT will be available at https://github.com/Gao-xiyuan/FCT. Yan Zhang 0108, Xiyuan Gao, Qingyan Duan, Jiaxu Leng, Xiao Pu 0002, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | DecomFormer: Decompose Self-Attention Via Fourier Transform for VHR Aerial Image Scene ClassificationabstractVery high-resolution (VHR) aerial image scene classification is an essential task for aerial image understanding. Although transformer-based models have demonstrated strong ability in natural image classification, transformer-based methods on VHR aerial image tasks are still lack of concern because the complexity of self-attention in the transformer grows quadratically with the image resolution. To address this issue, we decompose the self-attention via Fourier Transform and propose a novel Fourier self-attention (FSA) mechanism. Based on FSA, we design a highly efficient network named DecomFormer, which learns contextual relationships in the real part and imaginary part of the Fourier field, respectively. Theoretically, the DecomFormer reduces the complexity of the naive self-attention mechanism from O(n2) to O(nlog(n)). Universal experiments on public VHR aerial image classification benchmarks demonstrated the DecomFormer’s efficiency, especially on images with very high-resolution. Yan Zhang 0108, Xiyuan Gao, Xiao Pu 0002, Xinbo Gao 0001 |
ICASSP | 3 |
| 2023 | FCIR: Rethink Aerial Image Super Resolution with Fourier AnalysisabstractRecent years, deep-learning-based methods achieve remarkable improvements on the super-resolution (SR) task. However, recovering high-quality (HQ) texture from the low-quality (LQ) aerial image is still challenging due to the limited contextual modeling ability of current deep-learning methods as well as the sharp artificial texture of aerial images. In this paper, we rethink aerial image super resolution (AISR) task with the perspective of Fourier analysis. Firstly, we build the Fourier Global Convolution (FGC) inspired by the convolution theorem of the Fourier Transform to extract the shadow features. Then, following the Gabor Transform, a carefully designed oriented Texture Contextual Block (OTCB) is proposed to enhance the oriented texture representation. By stacking FGC and OTCB, we propose a simple but effective straight-forward network named Fourier Consistency Image Reconstruction Model (FCIR) to restore HQ aerial image. Moreover, we design a gradient consistency loss (GC Loss) to enhance the quality of reconstructed high-frequency details. Compared with very recent state-of-the-art super-resolution methods, experimental results demonstrate promising SR performance boosts from FCIR on 3 typical aerial image datasets. Yan Zhang 0108, Jianan Jiang, Xiao Pu 0002, Xinbo Gao 0001 |
ICASSP | 4 |
| 2023 | Lexical knowledge enhanced text matching via distilled word sense disambiguation
Xiao Pu 0002, Lin Yuan 0002, Jiaxu Leng, Tao Wu 0003, Xinbo Gao 0001 |
Knowl. Based Syst. | 1 |
| 2022 | PRO-Face: A Generic Framework for Privacy-preserving Recognizable Obfuscation of Face ImagesabstractA number of applications (e.g., video surveillance and authentication) rely on automated face recognition to guarantee functioning of secure services, and meanwhile, have to take into account the privacy of individuals exposed under camera systems. This is the so-called Privacy-Utility trade-off. However, most existing approaches to facial privacy protection focus on removing identifiable visual information from images, leaving protected face unrecognizable to machine, which sacrifice utility for privacy. To tackle the privacy-utility challenge, we propose a novel, generic, effective, yet lightweight framework for Privacy-preserving Recognizable Obfuscation of Face images (named as PRO-Face). The framework allows one to first process a face image using any preferred obfuscation, such as image blur, pixelate and face morphing. It then leverages a Siamese network to fuse the original image with its obfuscated form, generating the final protected image visually similar to the obfuscated one from human perception (for privacy) but still recognized as the original identity by machine (for utility). The framework supports various obfuscations for facial anonymization. The face recognition can be performed accurately not only across anonymized images but also between plain and anonymized ones, based on only pre-trained recognizers. Those feature the "generic" merit of the proposed framework. In-depth objective and subjective evaluations demonstrate the effectiveness of the proposed framework in both privacy protection and utility preservation under distinct scenarios. Our source code, models and any supplementary materials are made publicly available. Lin Yuan 0002, Linguo Liu, Xiao Pu 0002, Xinbo Gao 0001 |
ACM Multimedia | 3 |