VLDB 2026 Research / reviewers in the wild / expert
Xin Peng 0010
dblp:14/6370-10
· DBLP profile ↗
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0001-8642-8582ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-View Progressive Feature Filtering for Multi-View Graph Clustering in Remote SensingabstractMulti-view clustering of remote sensing data plays a vital role in Earth observation analysis. Recently, deep graph clustering methods based on contrastive learning have significantly improved feature representation capabilities. However, most existing approaches treat all views equally, neglecting the inherent uniqueness and heterogeneity across views, which often results in two major issues: 1) discriminative features from clustering-friendly views are underexplored; and 2) redundant or noisy information from less informative views can degrade the shared representation. To address these challenges, we propose a novel multi-view graph clustering framework termed CF-MVGC for remote sensing data, which dynamically preserves discriminative features and suppresses redundancy by assessing view affinity. Specifically, we employ a dual-stage representation learning strategy to extract both view-specific discriminative features and cross-view consistent representations. To further exploit and adaptively integrate complementary information across views, we design a progressive feature filtering model that dynamically evaluates view affinity using two novel metrics, i.e., view fidelity index (VFI) and view criticality index (VCI). Based on these assessments, the module adaptively modulates feature update and reset signals, reinforcing informative views while suppressing noisy or redundant ones. Views with high affinity receive strengthened update signals to retain valuable features, while those with low affinity are subjected to enhanced reset operations to eliminate noise and redundancy. The resulting high-quality, discriminative representations lead to improved clustering performance, establishing a positive feedback loop. Experimental results on four benchmark datasets demonstrate the effectiveness and superiority of CF-MVGC against its competitors. Bowen Liu 0020, Xin Peng 0010, Wenxuan Tu, Chengyao Wei, Xiangyan Tang, Jieren Cheng |
AAAI | 2 |
| 2026 | Clara: A Cross-Modal Learning Framework for Enhanced Vulnerability DetectionabstractSoftware vulnerability detection is crucial for ensuring the security of software systems, representing a significant and challenging task. Recently, some studies have integrated large language models and graph neural networks to extract code features from different modalities (code sequences and graphs) for vulnerability detection. Unfortunately, current solutions struggle to fully leverage the complementary knowledge between modalities, thereby undermining their effectiveness in practical applications. In this paper, we proposeClara, a novel cross-modal learning approach that integrates multi-modal information from both global and local perspectives for effective detection. Specifically, for local fusion, we design an information interaction module guided by prompts, which employs learnable prompts to enhance feature extraction through the interaction of information between modalities. For global fusion, we devise a Cross-attention Adaptive Fusion module that adaptively adjusts the fusion weights of embeddings from different modalities using attention mechanisms. Experimental results on two benchmark datasets demonstrate thatClaraachieves improvements of 21.37% and 11.86% in F1 score over state-of-the-art vulnerability detection methods, respectively. Xin Peng 0010, Shangwen Wang, Bo Lin 0011, Yihao Qin, Liqian Chen, Xiaoguang Mao |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2026 | Fault Localization from the Semantic Code Search PerspectiveabstractThe software development process is characterized by an iterative cycle of continuous functionality implementation and debugging, essential for the enhancement of software quality and adaptability to changing requirements. This process incorporates two isolatedly studied tasks: Code Search (CS), which retrieves reference code from a code corpus to aid in code implementation, and Fault Localization (FL), which identifies code entities responsible for bugs within the software project to boost software debugging. The basic observation of this study is that these two tasks exhibit similarities since they both address search problems. Notably, CS techniques have demonstrated greater effectiveness than FL ones, possibly because of the precise semantic details of the required code offered by natural language queries, which are not readily accessible to FL methods. Drawing inspiration from this, we hypothesize that a fault localizer could achieve greater proficiency if semantic information about the buggy methods were made available. Based on this idea, we propose \(\texttt{CosFL}\) , an FL approach that decomposes the FL task into two steps: query generation , which describes the functionality of the problematic code in natural language, and fault retrieval , which uses CS to find program elements semantically related to the query, allowing for finishing the FL task from a CS perspective. Specifically, to depict the buggy functionalities and generate high-quality queries, \(\texttt{CosFL}\) extensively harnesses the code analysis, semantic comprehension, text generation, and decision-making capabilities of LLMs. Moreover, to enhance the accuracy of CS, \(\texttt{CosFL}\) captures varying levels of context information and employs a multi-granularity CS strategy, which facilitates a more precise identification of buggy methods from a holistic view. The evaluation on 835 real bugs from 23 Java projects shows that \(\texttt{CosFL}\) successfully localizes 324 bugs within Top-1, which significantly outperforms the state-of-the-art approaches by 26.6%–57.3%. The ablation study and sensitivity analysis further validate the importance of different components and the robustness of \(\texttt{CosFL}\) across different backend models. Yihao Qin, Shangwen Wang, Yan Lei 0005, Zhuo Zhang 0007, Bo Lin 0011, Xin Peng 0010, Jun Ma 0015, Liqian Chen, Xiaoguang Mao |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2026 | AGCN: A Service Computing-Oriented Anchor-Guided Graph Clustering NetworkabstractDeep graph clustering aims to uncover latent structural patterns within user behavior graphs to effectively group users based on behavioral similarity, thereby enhancing the efficiency of personalized service recommendation. However, most existing generative clustering methods heavily rely on encoder–decoder architectures and their pretraining procedures, which not only increase training complexity but also result in insufficient robustness and unstable optimization. To address these limitations, this paper proposes an Anchor-Guided Graph Clustering Network (AGCN) that eliminates the need for pretraining. The proposed method introduces an anchor-guided learning strategy that automatically selects high-confidence anchor samples farthest from the distribution boundary and utilizes the k-means++ algorithm to generate pseudo-labels in the diffusion feature space. Furthermore, AGCN integrates an anchor expansion strategy and a commonality enhancement mechanism to significantly improve clustering discriminability and recommendation accuracy. Theoretical analysis and extensive experiments on multiple real-world datasets demonstrate that the proposed method outperforms state-of-the-art approaches in terms of both clustering accuracy and stability. Jieren Cheng, Faqiang Zeng, Jimei Li, Ling Liu 0001, Naixue Xiong, Xin Peng 0010 |
IEEE Trans. Serv. Comput. | 6 |
| 2025 | Federated Graph-Level Clustering NetworkabstractFederated graph learning (FGL), which excels in analyzing non-IID graphs as well as protecting data privacy, has recently emerged as a hot topic. Existing FGL methods usually train the client model using labeled data and then collaboratively learn a global model without sharing their local graph data. However, in real-world scenarios, the lack of data annotations impedes the negotiation of multi-source information at the server, leading to sub-optimal feedback to the clients. To address this issue, we propose a novel unsupervised learning framework called Federated Graph-level Clustering Network (FedGCN), which collects the topology-oriented features of non-IID graphs from clients to generate global consensus representations through multi-source clustering structure sharing. Specifically, in the client, we first preserve the prototype features of each cluster from the structure-oriented embedding through clustering and then upload the learned multiple prototypes that are hard to be reconstructed into the raw graph data. In the server, we generate consensus prototypes from multiple condensed structure-oriented signals through Gaussian estimation, which are subsequently transferred to each client to promote the great encoding capacity of the local model for better clustering. Extensive experiments across multiple non-IID graph datasets have demonstrated the effectiveness and superiority of FedGCN against its competitors. Jingxin Liu 0006, Jieren Cheng, Renda Han, Wenxuan Tu, Xin Peng 0010 |
AAAI | 6 |
| 2025 | Let the Code Speak: Incorporating Program Dynamic State for Better Method-Level Fault LocalizationabstractFault localization (FL) is a critical but time-consuming part of software debugging. With the improvement of the Large Language Models (LLMs) in their code capabilities, the increasing demand for automated software development has encouraged more research on building LLM-based Fault Localization (LLMFL) systems. However, existing LLMFL techniques are typically restricted to predicting bug locations by analyzing static code, while overlooking crucial dynamic program state of the software. This lack of context makes LLMs prone to generating "hallucinations", incorrectly identifying bug-free code as suspicious. To address this, this paper introduces PingFL, the LLMFL system that incorporates program dynamic information for more accurate automatic fault localization. PingFL comprises a Fault Localization (FL) agent and a Print Debugging (PD) agent. The FL agent is tasked with understanding the root cause through a set of callable tools. When the FL agent nominates a location as suspicious, it would entrust the PD agent to verify the suspected issue through multiple rounds of print debugging. In particular, these two agents communicate efficiently by conveying the textual thought generated by the LLM. The evaluation on 812 real-world bugs from the Defects4J benchmark shows that PingFL can localize 450 bugs within Top-1, which significantly outperforms other LLM-based approaches by 41% to 122%. A deeper dive into PingFL’s performance reveals that it exhibits specific FL strategies and tool usage patterns even without explicit instructions. Finally, PingFL proves to be cost-effective, spending an average of $0.23 and 104.62 seconds per bug, with the print debugging mechanism accounting for only $0.07 and 48.14 seconds. Yihao Qin, Shangwen Wang, Bo Lin 0011, Xin Peng 0010, Sheng Ouyang, Liqian Chen, Xiaoguang Mao |
ASE | 4 |
| 2025 | Multi-view Graph Clustering with Dual Structure Awareness for Remote Sensing DataabstractMulti-view clustering plays a pivotal role in remote sensing image analysis, where graph neural network-based methods have demonstrated remarkable potential by modeling data as graphs. However, existing efforts, which construct remote sensing graphs using fixed rules (e.g., K-nearest neighbors), inevitably introduce noisy edges and increase the risk of heterogeneous information diffusion, leading to inferior clustering performance. Although recent works attempt to address this issue by refining the structure, they are designed for single-view data and struggle to extend to multi-view scenarios. To bridge this gap, we propose a dual structure awareness multi-view graph clustering method named DSMVGC, which generates two distinct structures for each view through explicit and implicit perspectives. Specifically, in our method, the learning processes of structure refinement and clustering are alternately optimized to mutually enhance each other. On one hand, the explicit structure updates the topology based on inter-cluster relationships, while the implicit structure captures latent relationships not covered by the explicit structure through adversarial learning. On the other hand, the refined structures not only facilitate homogeneous message passing but also serve as prior knowledge to guide the contrastive loss, thereby enhancing the discriminability of representations for accurate clustering. Extensive experiments on five multi-view remote sensing datasets validate the effectiveness of DSMVGC. Xin Peng 0010, Bowen Liu 0020, Renxiang Guan, Wenxuan Tu |
ACM Multimedia | 1 |
| 2025 | Keep It Simple: Self-Adaptive Code Graph Simplification for Accurate Vulnerability DetectionabstractSoftware vulnerability detection is crucial for high-quality software development. Recently, some studies utilizing Graph Neural Networks (GNNs) to learn the graph representation of code in vulnerability detection tasks have achieved remarkable success. However, existing graph-based approaches mainly face two limitations that prevent them from generalizing well to large code graphs: (1) the interference of noise information in the code graph; (2) the difficulty in capturing long-distance dependencies within the graph. To mitigate these problems, we propose a novel vulnerability detection method,ANGEL, whose novelty mainly embodies the hierarchical graph refinement and context-aware graph representation learning. The former hierarchically filters redundant information in the code graph, thereby reducing the size of the graph, while the latter collaboratively employs the Graph Transformer and GNN to learn code graph representations from both the global and local perspectives, thus capturing long-distance dependencies. Extensive experiments demonstrate promising results on three widely used benchmark datasets: our method significantly outperforms several other baselines in terms of the accuracy and F1 score. Particularly, in large code graphs,ANGELachieves an improvement in accuracy of 34.27%-161.93% compared to the state-of-the-art method, AMPLE. Such results demonstrate the effectiveness ofANGELin vulnerability detection tasks. Xin Peng 0010, Shangwen Wang, Yihao Qin, Bo Lin 0011, Liqian Chen, Jieren Cheng, Xiaoguang Mao |
IEEE Trans. Software Eng. | 1 |
| 2024 | Attribute-Missing Graph Clustering NetworkabstractDeep clustering with attribute-missing graphs, where only a subset of nodes possesses complete attributes while those of others are missing, is an important yet challenging topic in various practical applications. It has become a prevalent learning paradigm in existing studies to perform data imputation first and subsequently conduct clustering using the imputed information. However, these ``two-stage" methods disconnect the clustering and imputation processes, preventing the model from effectively learning clustering-friendly graph embedding. Furthermore, they are not tailored for clustering tasks, leading to inferior clustering results. To solve these issues, we propose a novel Attribute-Missing Graph Clustering (AMGC) method to alternately promote clustering and imputation in a unified framework, where we iteratively produce the clustering-enhanced nearest neighbor information to conduct the data imputation process and utilize the imputed information to implicitly refine the clustering distribution through model optimization. Specifically, in the imputation step, we take the learned clustering information as imputation prompts to help each attribute-missing sample gather highly correlated features within its clusters for data completion, such that the intra-class compactness can be improved. Moreover, to support reliable clustering, we maximize inter-class separability by conducting cost-efficient dual non-contrastive learning over the imputed latent features, which in turn promotes greater graph encoding capability for clustering sub-network. Extensive experiments on five datasets have verified the superiority of AMGC against competitors. Wenxuan Tu, Renxiang Guan, Sihang Zhou 0001, Chuan Ma 0001, Xin Peng 0010, Zhiping Cai, Zhe Liu 0001, Jieren Cheng, Xinwang Liu 0002 |
AAAI | 5 |
| 2024 | RARE: Robust Masked Graph AutoencoderabstractMasked graph autoencoder (MGAE) has emerged as a promising self-supervised graph pre-training (SGP) paradigm due to its simplicity and effectiveness. However, existing efforts perform the mask-then-reconstruct operation in the raw data space as is done in computer vision (CV) and natural language processing (NLP) areas, while neglecting the important non-Euclidean property of graph data. As a result, the highly unstable local structures largely increase the uncertainty in inferring masked data and decrease the reliability of the exploited self-supervision signals, leading to inferior representations for downstream evaluations. To address this issue, we propose a novel SGP method termed Robust mAsked gRaph autoEncoder (RARE) to improve the certainty in inferring masked data and the reliability of the self-supervision mechanism by further masking and reconstructing node samples in the high-order latent feature space. Through both theoretical and empirical analyses, we have discovered that performing a joint mask-then-reconstruct strategy in both latent feature and raw data spaces could yield improved stability and performance. To this end, we elaborately design a masked latent feature completion scheme, which predicts latent features of masked nodes under the guidance of high-order sample correlations that are hard to be observed from the raw data perspective. Specifically, we first adopt a latent feature predictor to predict the masked latent features from the visible ones. Next, we encode the raw data of masked samples with a momentum graph encoder and subsequently employ the resulting representations to improve the predicted results through latent feature matching. Extensive experiments on seventeen datasets have demonstrated the effectiveness and robustness of RARE against state-of-the-art (SOTA) competitors across three downstream tasks. Our source code is available athttps://github.com/WxTu/RARE. Wenxuan Tu, Qing Liao 0001, Sihang Zhou 0001, Xin Peng 0010, Chuan Ma 0001, Zhe Liu 0001, Xinwang Liu 0002, Zhiping Cai, Kunlun He |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Dual Contrastive Learning Network for Graph ClusteringabstractGraph representation is an important part of graph clustering. Recently, contrastive learning, which maximizes the mutual information between augmented graph views that share the same semantics, has become a popular and powerful paradigm for graph representation. However, in the process of patch contrasting, existing literature tends to learn all features into similar variables, i.e., representation collapse, leading to less discriminative graph representations. To tackle this problem, we propose a novel self-supervised learning method called dual contrastive learning network (DCLN), which aims to reduce the redundant information of learned latent variables in a dual manner. Specifically, the dual curriculum contrastive module (DCCM) is proposed, which approximates the node similarity matrix and feature similarity matrix to a high-order adjacency matrix and an identity matrix, respectively. By doing this, the informative information in high-order neighbors could be well collected and preserved while the irrelevant redundant features among representations could be eliminated, hence improving the discriminative capacity of the graph representation. Moreover, to alleviate the problem of sample imbalance during the contrastive process, we design a curriculum learning strategy, which enables the network to simultaneously learn reliable information from two levels. Extensive experiments on six benchmark datasets have demonstrated the effectiveness and superiority of the proposed algorithm compared with state-of-the-art methods. Xin Peng 0010, Jieren Cheng, Xiangyan Tang, Jingxin Liu 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | HSNet: An Intelligent Hierarchical Semantic-Aware Network System for Real-Time Semantic SegmentationabstractSemantic segmentation, which aims to accurately identify each pixel, is a meaningful and challenging task. Recently, we witness a strong tendency to improve model efficiency in low-computing applications. However, most real-time methods ignore hierarchical features and context information to improve efficiency, leading to a decrease in the accuracy of semantic segmentation. To this end, we propose a novel system named hierarchical semantic-aware network (HSNet) to refine multilevel context information. HSNet mainly has the following two core modules: 1) hierarchical feature refinement module (HFRM) and 2) cross-scale pyramid fusion module (CPFM). By aggregating hierarchical feature maps, the proposed HFRM learns multilevel feature representation to recover spatial details. Afterward, the dual attention mechanism is developed to refine features from both channel and spatial levels, thereby alleviating the multilevel semantic gap. Meanwhile, the CPFM, which fuses local and global context information in a cross-scale manner, is proposed to enrich semantic information to improve accuracy. Furthermore, HSNet is carefully designed to improve the efficiency of the model by reusing shallow features and reducing channel capacity. Extensive experiments show that our method is effective and superior in segmentation accuracy and inference speed compared with state-of-the-art methods. Xin Peng 0010, Jieren Cheng, Xiangyan Tang, Ziqi Deng, Wenxuan Tu, Naixue Xiong |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2022 | MIFNet: A lightweight multiscale information fusion networkabstractSemantic segmentation technique plays a crucial role in Internet of Things applications, such as industrial robotics and self-driving. Recently deep learning approaches have boosted semantic segmentation accuracy greatly. However, their comprehensive performance in terms of accuracy and efficiency is still far from satisfactory. We observe that (1) accuracy-oriented methods rely on numerous convolution layers and sophisticated architectures, which result in heavy computational complexity and usually take a long time for inference; (2) efficiency-oriented methods fail to capture the multiscale context information for discriminative representations during the feature fusion process, thus leading to suboptimal performance. Previous semantic segmentation approaches fail to address these two challenges simultaneously. To tackle the dilemma of precise segmentation and efficient inference, we propose a novel lightweight Multiscale Information Fusion Network (MIFNet). Specifically, the proposed MIFNet mainly consists of two core components, that is, Pyramid Refinement Connection Module (PRCM) and Lightweight Information Fusion Module (LIFM). The PRCM exploits skip learning to establish dependency between different stages. Meanwhile, the pyramid attention mechanism (PAM) in PRCM, which adjusts the weight of hybrid pyramid attention vector to refine spatial features of low-level, is developed to alleviate the semantic gap. Moreover, the LIFM is designed to detect objects at multiple scales from the global-local perspective. In LIFM, the proposed multiscale dense concatenation (MDC) adopts various dilated convolution to extract multiscale local context information. Extensive experimental results on benchmarks data sets demonstrate the significantly better performance of the proposed MIFNet compared with most existing state-of-the-art methods. Jieren Cheng, Xin Peng 0010, Xiangyan Tang, Wenxuan Tu, Wenhang Xu |
Int. J. Intell. Syst. | 2 |