VLDB 2026 Research / reviewers in the wild / expert
Xiaodan Xu
dblp:65/8840
· DBLP profile ↗
11ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Abundant Modalities Offer More Nutrients: Multi-Modal-Based Function-Level Vulnerability DetectionabstractSoftware vulnerabilities are weaknesses in software systems that can lead to significant cybersecurity risks. Recently, several deep learning (DL)-based approaches have been proposed to detect vulnerabilities at the function level. These approaches typically utilize one or a few different modalities (e.g., text representation and graph-based representation) of the function, and have shown promising performance. However, existing studies have not fully leveraged diverse modalities, particularly those that use images to represent functions for vulnerability detection. These approaches often fail to make sufficient use of the important graph structure underlying the images. In this article, we propose MVulD+, a multi-modal-based function-level vulnerability detection approach, which fuses multi-modal features of the function (i.e., text representation, graph representation, and image representation) to detect vulnerabilities. Specifically, MVulD+ leverages a pre-trained model (i.e., UniXcoder) to capture the semantic information of the textual source code, uses a graph neural network to extract graph representations, and employs computer vision techniques to obtain image representations while preserving the graph structure of the function. To investigate the effectiveness of MVulD+, we conduct a large-scale experiment by comparing our approach with nine state-of-the-art baselines. Experimental results demonstrate that MVulD+ improves the DL-based baselines by 24.3–125.7%, 5.2–31.4%, 40.6–192.2%, and 22.3–186.9% in terms of F1-score, Accuracy, Precision, and PR-AUC, respectively. Chao Ni 0001, Xinrui Li 0004, Xiaodan Xu |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2025 | What You See is What You Get: Attention-Based Self-Guided Automatic Unit Test GenerationabstractSoftware defects heavily affect software's function-alities and may cause huge losses. Recently, many AI-based approaches have been proposed to detect defects, which can be divided into two categories: software defect prediction and automatic unit test generation. While these approaches have made great progress in software defect detection, they still have several limitations in practical application, including the low confidence of prediction models and the inefficiency of unit testing models. To address these limitations, we propose a WYSIWYG (i.e., What You See Is What You Get) approach: Attention-based Self-guided Automatic Unit Test GenERation (AUGER), which contains two stages: defect detection and error triggering. In the former stage, AUGER first detects the proneness of defects. Then, in the latter stage, it guides to generate unit tests for triggering such an error with the help of critical information obtained by the former stage. To evaluate the effectiveness of AUGER, we conduct a large-scale experiment by comparing with the state-of-the-art (SOTA) approaches on the widely used datasets (i.e., Bears, Bugs.jar, and Defects4J). AUGER makes great improvements by 4.7% to 35.3% and 17.7% to 40.4% in terms of Fl-score and Precision in defect detection, and can trigger 23 to 84 more errors than SOTAs in unit test generation. Besides, we also conduct a further study to verify the generalization in practical usage by collecting a new dataset from real-world projects. Chao Ni 0001, Xiaodan Xu, Xiaohu Yang 0001 |
ICSE | 3 |
| 2025 | Sembug: Detecting Logic Bugs in Dbms Through Generating Semantic-Aware Non-Optimizing QueryabstractLogic bugs, which cause Database Management Systems (DBMSs) to return incorrect results, are challenging to detect due to the absence of explicit signs such as system crashes. The majority of these bugs originate from the query optimizer and are commonly referred to as optimization bugs. Many approaches have been proposed for detecting logic bugs, which can be divided into two groups. The first group aims to detect the optimization bugs but only focuses on those with incorrect results cardinality, neglecting to check semantic correctness and consequently limiting the detection of bugs in advanced DBMS features. For the second group, though it can verify the correctness of the results for both their cardinality and semantics, it is ineffective in handling optimization bugs, which restricts its practical usage effectiveness. In this paper, we propose Semantic-aware Non-Optimizing Query (SemBug), a novel approach for logic bug detection in DBMSs. SemBug focuses on optimization bugs by transforming the queries that can be highly optimized by DBMS into equivalent but less optimized ones. Additionally, SemBug integrates semantic analysis technology, enabling it to identify semantic logic bugs and support testing advanced DBMS features. Any discrepancy in cardinality or content between the original and transformed queries indicates a logic bug. To investigate the effectiveness of SemBug, we conduct a large-scale experiment on five widelyused DBMS systems (i.e., MySQL, TiDB, MariaDB, SQLite, and PostgreSQL) and compare it with three state-of-the-art (SOTA) approaches (i.e., Pinolo, TLP, and NoREC). The experimental results indicate that SemBug outperforms three SOTAs. Over 24 hours, SemBug found 34 unique logic bugs, which are 19, 14, and 13 more bugs than each of the three SOTAs, marking an improvement of$126 \%, 70 \%$, and 61 % respectively. As of the time of paper submission, SemBug has uncovered 37 unique logic bugs, of which 29 have been verified by developers, and 11 have been fixed. SemBug helps developers identify these bugs, providing insights into such inconsistencies and assisting in resolving them. Shiyang Ye, Chao Ni 0001, Qianqian Pang, Xinrui Li 0004, Xiaodan Xu |
ICPC | 6 |
| 2025 | Adaptive Feature Fusion Enhanced Cascade Pointer Network for Chinese Relation Extraction
Jiaqing Shi, Xiayan Ji, Lina Chen, Hong Gao 0001, Xiaodan Xu |
WASA (1) | 5 |
| 2025 | Distinguishing LLM-Generated from Human-Written Code by Contrastive LearningabstractLarge language models (LLMs), such as ChatGPT released by OpenAI, have attracted significant attention from both industry and academia due to their demonstrated ability to generate high-quality content for various tasks. Despite the impressive capabilities of LLMs, there are growing concerns regarding their potential risks in various fields, such as news, education, and software engineering. Recently, several commercial and open source LLM-generated content detectors have been proposed, which, however, are primarily designed for detecting natural language content without considering the specific characteristics of program code. This article aims to fill this gap by proposing a novel ChatGPT-generated code detector, CodeGPTSensor, based on a contrastive learning framework and a semantic encoder built with UniXcoder. To assess the effectiveness of CodeGPTSensor on differentiating ChatGPT-generated code from human-written code, we first curate a large-scale Human and Machine comparison Corpus (HMCorp), which includes 550k pairs of human-written and ChatGPT-generated code (i.e., 288k Python code pairs and 222k Java code pairs). Based on the HMCorp dataset, our qualitative and quantitative analysis of the characteristics of ChatGPT-generated code reveals the challenge and opportunity of distinguishing ChatGPT-generated code from human-written code with their representative features. Our experimental results indicate that CodeGPTSensor can effectively identify ChatGPT-generated code, outperforming all selected baselines. Xiaodan Xu, Chao Ni 0001, Xinrong Guo, Shaoxuan Liu, Kui Liu 0001, Xiaohu Yang 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2024 | Macroscopic Emission Modeling of Urban Traffic Using Probe Vehicle Data: A Machine Learning ApproachabstractUrban congestions cause inefficient movement of vehicles and exacerbate greenhouse gas emissions and urban air pollution. Macroscopic emission fundamental diagram (eMFD) captures an orderly relationship among emission and aggregated traffic variables at the network level, allowing for real-time monitoring of region-wide emissions and optimal allocation of travel demand to existing networks, reducing urban congestion and associated emissions. However, empirically derived eMFD models are sparse due to historical data limitation. Leveraging a large-scale and granular traffic and emission data derived from probe vehicles, this study is the first to apply machine learning methods to predict the network-wide emission rate to traffic relationship in U.S. urban areas at a large scale. The analysis framework and insights developed in this work generate data-driven eMFDs and a deeper understanding of their location dependence on network, infrastructure, land use, and vehicle characteristics, enabling transportation authorities to measure carbon emissions from urban transport of given travel demand and optimize location-specific traffic management and planning decisions to mitigate network-wide emissions. Mohammed Adlouni, Ling Jin 0001, Xiaodan Xu, Anna Spurlock, Alina Lazar, Kaveh Farokhi Sadabadi, Mahyar Amirgholy, Mona Asudegi |
IEEE Big Data | 3 |
| 2023 | Leveraging Probe Data and Machine Learning to Derive and Interpret Macroscopic Fundamental Diagrams Across U.S. CitiesabstractMacroscopic fundamental diagram (MFD) captures an orderly relationship among traffic flow, density, and speed at the network level. Understanding network-wide traffic through MFDs can optimally allocate demand to existing networks, improving performance by maximizing network production and avoiding congestion. However, due to historical data limitations, empirically derived MFD models are sparse in the literature, especially for the U.S. cities. Leveraging a large-scale and granular census-tract-level flow and density derived from vehicle probe data, this research is the first to develop a machine learning approach to both derive MFD models and interpret their underlying difference among urban networks across the entire United States. Among the four machine learning methods tested here XGBoost is found to deliver the best performance to predict the network traffic flow for given vehicular density and location attributes. Interaction Shapley Additive explanation (SHAP) values are used to interpret the factors, such as land use, transportation infrastructure, and network topology, that influence the flow-density relationships among locations. The analysis framework developed in this work can generate datadriven MFDs and a deeper understanding of their shape dependence on network, infrastructure, and land use characteristics, which can be used by transportation authorities to derive and optimize location-specific MFDs facilitating more informed management and planning decisions at the network level. Ling Jin 0001, Xiaodan Xu, Kaveh Farokhi Sadabadi, Alina Lazar, Duleep Rathgamage Don, Zachary Needell, Anna Spurlock, Mahyar Amirgholy, Mona Asudegi |
IEEE Big Data | 2 |
| 2023 | Function-Level Vulnerability Detection Through Fusing Multi-Modal KnowledgeabstractSoftware vulnerabilities damage the functionality of software systems. Recently, many deep learning-based approaches have been proposed to detect vulnerabilities at the function level by using one or a few different modalities (e.g., text representation, graph-based representation) of the function and have achieved promising performance. However, some of these existing studies have not completely leveraged these diverse modalities, particularly the underutilized image modality, and the others using images to represent functions for vulnerability detection have not made adequate use of the significant graph structure underlying the images. In this paper, we propose MVulD, a multi-modal-based function-level vulnerability detection approach, which utilizes multi-modal features of the function (i.e., text representation, graph representation, and image representation) to detect vulnerabilities. Specifically, MVulD utilizes a pre-trained model (i.e., UniXcoder) to learn the semantic information of the textual source code, employs the graph neural network to distill graph-based representation, and makes use of computer vision techniques to obtain the image representation while retaining the graph structure of the function. We conducted a large-scale experiment on 25,816 functions. The experimental results show that MVulD improves four state-of-the-art baselines by 30.8%-81.3%, 12.8%-27.4%, 48.8%-115%, and 22.9%-141% in terms of F1-score, Accuracy, Precision, and PR-AUC respectively. Chao Ni 0001, Xinrong Guo, Xiaodan Xu, Xiaohu Yang 0001 |
ASE | 4 |
| 2023 | Boosting Just-in-Time Defect Prediction with Specific Features of C/C++ Programming Languages in Code ChangesabstractJust-in-time (JIT) defect prediction can identify changes as defect-inducing ones or clean ones and many approaches are proposed based on several programming language-independent change-level features. However, different programming languages have different characteristics and consequently may affect the quality of software projects. Meanwhile, the C programming language, one of the most popular ones, is widely used to develop foundation applications (i.e., operating system, database, compiler, etc.) in IT companies and its change-level characteristics on project quality have not been fully investigated. Additionally, whether open-source C projects have similar important features to commercial projects has not been studied much.To address the aforementioned limitations, in this paper, we investigate the impacts of programming language-specific features on the state-of-the-art JIT defect identification approach in an industrial setting. We collect and label the top-10 most starred C projects (i.e., 329,021 commits) on GitHub and 8 C projects in an ICT company (i.e., 12,983 commits). We also propose nine C-specific change-level features and focus our investigations on both open-source C projects on GitHub and C projects at the ICT company considering three aspects: (1) The effectiveness of C-specific change-level features in improving the performance of identification of defect-inducing changes, (2) The importance of features in the identification of defect-inducing changes between open-source C projects and commercial C projects, and (3) The effectiveness of combining language-independent features and C-specific features in a real-life setting at the ICT company. Chao Ni 0001, Xiaodan Xu, David Lo 0001 |
MSR | 2 |
| 2023 | Anomaly Detection With Representative NeighborsabstractIdentifying anomalies from data has attracted increasing attention in recent years due to its broad range of potential applications. Although many efforts have been made for anomaly detection, how to effectively handle high-dimensional data and how to exactly explore neighborhood information, a fundamental issue in anomaly detection, have not yet received sufficient concerns. To circumvent these challenges, in this article, we propose an effective anomaly detection method with representative neighbors for high-dimensional data. Specifically, it projects the high-dimensional data into a low-dimensional space via a sparse operation and explores representative neighbors with a self-representation learning technique. The neighborhood information is then transformed into similarity relations, making the data converge or disperse. Eventually, anomalies are discriminated by a tailored graph clustering technique, which can effectively reveal structural information of the data. Extensive experiments were conducted on ten public real-world datasets with 11 popular anomaly detection algorithms. The results show that the proposed method has encouraging and promising performance compared to the state-of-the-art anomaly detection algorithms. Huawen Liu, Xiaodan Xu, Enhui Li, Shichao Zhang 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | DISK: Domain-constrained Instance Sketch for Math Word Problem GenerationabstractA math word problem (MWP) is a coherent narrative which reflects the underlying logic of math equations. Successful MWP generation can automate the writing of mathematics questions. Previous methods mainly generate MWP text based on inflexible pre-defined templates. In this paper, we propose a neural model for generating MWP text from math equations. Firstly, we incorporate a matching model conditioned on the domain knowledge to retrieve a MWP instance which is most consistent with the ground-truth, where the domain is a latent variable extracted with a domain summarizer. Secondly, by constructing a Quantity Cell Graph (QCG) from the retrieved MWP instance and reasoning over it, we improve the model’s comprehension of real-world scenarios and derive a domain-constrained instance sketch to guide the generation. Besides, the QCG also interacts with the equation encoder to enhance the alignment between math tokens (e.g., quantities and variables) and MWP text. Experiments and empirical analysis on educational MWP set show that our model achieves impressive performance in both automatic evaluation metrics and human evaluation metrics. Tianyang Cao, Shuang Zeng, Xiaodan Xu, Mairgup Mansur, Baobao Chang |
COLING | 3 |