VLDB 2026 Research / reviewers in the wild / expert
Xiaomeng Wang 0001
dblp:35/6677-1
· DBLP profile ↗
9ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-7591-0127ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Multi-Feature Fusion Approach for Design Pattern Detection Based on Graph Neural NetworksabstractDesign patterns are powerful tools that provide standardized solutions to common problems in the software design process. Detecting design patterns in software systems can greatly facilitate software maintenance and code understanding, but manual detection is a complex and challenging task. In recent years, with the development of deep learning techniques, researchers have proposed automatic design pattern detection methods that utilize code features and machine learning classifiers. However, these approaches typically focus on either code semantic features or structural features, lacking the integration of both. To address this issue, in this paper, we propose a novel design pattern detection method with multi-feature fusion, DPDMF2, which effectively fuses code semantic features and software network structural features for design pattern detection. Specifically, we first represent software as a CSN (Class-level Software Network) to capture its topology, including classes, their relationships, and the types of these relationships. Concurrently, we utilize code representation learning techniques to obtain semantic features of the code, which are then used as node attribute features within the network. Finally, we utilize graph neural networks (GNNs) to fuse these two types of features, generating a vector representation that encapsulates both semantic and structural information for design pattern detection task. Experimental results on public datasets show that DPDMF2achieves 81.7% precision and 82.1% recall, validating the overall effectiveness of our method. Nong Zou, Junxiang Zhang, Xiaomeng Wang 0001, Tao Jia 0001 |
CSCWD | 4 |
| 2025 | Towards a performance characteristic curve for model evaluation: An application in information diffusion prediction
Wenjin Xie, Xiaomeng Wang 0001, Radoslaw Michalski, Tao Jia 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Effective Vulnerability Detection over Code Token Graph: A GCN with Score Gate Based ApproachabstractIn modern society, software systems are integral to various aspects of life. Finding an efficient vulnerability identification approach is crucial for ensuring security and preventing malicious attacks. In recent years, many deep learning-based methods have shown outstanding performance in the vulnerability detection task. However, these methods still have limitations. Some methods consider code input as token sequences and apply architectures typically used in natural language processing. They fail to utilize the structural information from various code components' interactions, which limits these models' performance. Other methods based on graph neural networks, although better at learning structural information, treat each node equally and fail to emphasize key elements. To overcome these limitations, we propose CTGGSG, an effective vulnerability detection approach over Code Token Graph (CTG) based on GCN with score gate. In our model, we use PLE-CG-SE module to represent the source code samples as the CTGs, effectively utilizing the high-quality feature representation of PL-PLM (Pre-trained Language Model based on Program Languages) and retaining structural information from the source code. During the graph learning process, we combine GCN convolution and score gate mechanism to make the model focus more on the key nodes within the graph and increase the receptive field of the nodes. To comprehensively evaluate the performance and scalability of our model, we conducted experiments on two real-world datasets: CodeX Glue,which contains balanced sample labels, and Re-veal, which contains imbalanced sample labels. These datasets contain 27,318 and 22,734 function-level samples, respectively, derived from large-scale, popular real-world projects. Compared to existing advanced vulnerability detection methods, our model achieved state-of-the-art performance overall. Nong Zou, Junxiang Zhang, Xiaomeng Wang 0001, Lai Hong, Tao Jia 0001 |
APSEC | 4 |
| 2024 | Design Pattern Representation and Detection Based on Heterogeneous Information Network
Tao Lu 0018, Xiaomeng Wang 0001, Tao Jia 0001 |
ICSR | 2 |
| 2022 | CCasGNN: Collaborative Cascade Prediction Based on Graph Neural NetworksabstractCascade prediction aims at modeling information diffusion in the network. Most previous methods concentrate on mining either structural or sequential features from the network and the propagation path. Recent efforts devoted to combining network structure and sequence features by graph neural networks and recurrent neural networks. Nevertheless, the limitation of spectral or spatial methods restricts the improvement of prediction performance. Moreover, recurrent neural networks are time-consuming and computation-expensive, which causes the inefficiency of prediction. Here, we propose a novel method CCasGNN considering the individual profile, structural features, and sequence information. The method benefits from using a collaborative framework of GAT and GCN and stacking positional encoding into the layers of graph neural networks, which is different from all existing ones and demonstrates good performance. The experiments conducted on two real-world datasets confirm that our method significantly improves the prediction accuracy compared to state-of-the-art approaches. Whats more, the ablation study investigates the contribution of each component in our method. Xiaomeng Wang 0001, Tao Jia 0001 |
CSCWD | 2 |
| 2022 | Independent Asymmetric Embedding for Information Diffusion Prediction on Social NetworksabstractThe prediction for information diffusion on social networks has great practical significance in marketing and public opinion control. It aims to predict the individuals who will potentially repost the message on the social network. One type of method is based on demographics, complex networks, and other prior knowledge to establish an interpretable model to simulate and predict the propagation process, while the other type of method is completely data-driven and maps the nodes to a latent space for propagation prediction. Existing latent space design and embedding methods lack consideration for the intervention among users. In this paper, we propose an independent asymmetric embedding method to embed each individual into one latent influence space and multiple latent susceptibility spaces. Based on the similarity between information diffusion and heat diffusion phenomenon, the heat diffusion kernel is exploited in our model and establishes the embedding rules. Furthermore, our method captures the co-occurrence regulation of user combinations in cascades to improve the calculating effectiveness. The results of extensive experiments conducted on real-world datasets verify both the predictive accuracy and cost-effectiveness of our approach. Wenjin Xie, Xiaomeng Wang 0001, Tao Jia 0001 |
CSCWD | 2 |
| 2022 | Author Name Disambiguation via Heterogeneous Network Embedding from Structural and Semantic PerspectivesabstractName ambiguity is common in academic digital libraries, such as multiple authors having the same name. This creates challenges for academic data management and analysis, thus name disambiguation becomes necessary. The procedure of name disambiguation is to divide publications with the same name into different groups, each group belonging to a unique author. A large amount of attribute information in publications makes traditional methods fall into the quagmire of feature selection. These methods always select attributes artificially and equally, which usually causes a negative impact on accuracy. The proposed method is mainly based on representation learning for heterogeneous networks and clustering and exploits the self-attention technology to solve the problem. The presentation of publications is a synthesis of structural and semantic representations. The structural representation is obtained by meta-path-based sampling and a skip-gram-based embedding method, and meta-path level attention is introduced to automatically learn the weight of each feature. The semantic representation is generated using NLP tools. Our proposal performs better in terms of name disambiguation accuracy compared with baselines and the ablation experiments demonstrate the improvement by feature selection and the meta-path level attention in our method. The experimental results show the superiority of our new method for capturing the most attributes from publications and reducing the impact of redundant information. Wenjin Xie, Xiaomeng Wang 0001, Tao Jia 0001 |
ICTAI | 3 |
| 2022 | CasSeqGCN: Combining network structure and temporal sequence to predict information cascadesabstractOne important task in the study of information cascade is to predict the future recipients of a message given its past spreading trajectory. While the network structure serves as the backbone of the spreading, an accurate prediction can hardly be made without the knowledge of the dynamics on the network. The temporal information in the spreading sequence captures many hidden features, but predictions based on sequence alone have their limitations. Recent efforts start to explore the possibility of combining both the network structure and the temporal feature. Here, we propose a new end-to-end prediction method CasSeqGCN in which the structure and temporal feature are simultaneously taken into account. A cascade is divided into multiple snapshots which record the network topology and the state of nodes. The graph convolutional network (GCN) is used to learn the representation of a snapshot. A novel aggregation method based on dynamic routing is proposed to aggregate node representation and the long short-term memory (LSTM) model is used to extract temporal information. CasSeqGCN predicts the future cascade size more accurately compared with other state-of-art baseline methods. The ablation study demonstrates that the improvement mainly comes from the design of the input and the GCN layer. We explicitly design an experiment to show the quality of the cascade representation learned by our approach is better than other methods. Our work proposes a new approach to combine the structural and temporal features, which not only gives a useful baseline model for future studies of cascade prediction, but also brings new insights on a wide collection of problems related with dynamics on and of the network. Xiaomeng Wang 0001, Yijun Ran, Radoslaw Michalski, Tao Jia 0001 |
Expert Syst. Appl. | 2 |
| 2020 | From Syntactic Structure to Semantic Relationship: Hypernym Extraction from Definitions by Recurrent Neural Networks Using the Part of Speech Information
Yixin Tan, Xiaomeng Wang 0001, Tao Jia 0001 |
ISWC (1) | 2 |