VLDB 2026 Research / reviewers in the wild / expert
Di Liu 0021
dblp:15/1777-21
· DBLP profile ↗
7ranked-venue papers
5as first author
6since 2021 · last 2025
0009-0007-2761-9427ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards understanding the security issues of Python programsabstractPython programming language has witnessed a steady increase in popularity over the past few decades.Renowned for its conciseness and readability, as well as its ease of learning and use, Python is widespread adoption has inevitably exposed it to a higher likelihood of encountering issues.Given that numerous code modifications exhibit repetitive and analogous patterns, an extensive examination of Python code-fixing patterns becomes imperative.Among these patterns, security-related issues hold significant importance due to their heightened risks and potential for substantial impact.Consequently, conducting research on security-related matters assumes utmost significance.In this paper, we conduct a thorough investigation to gain insights into the security issues prevalent in Python programs.Our approach involves collecting 413 popular open-source Python projects from GitHub and identifying 9,782 bug reports related to security concerns and their corresponding bug fixes.We employ automated clustering and manual summarization techniques, ultimately classifying them into 12 distinct categories, with six categories being of notable prevalence.We analyze the bug reports and commits within each high-frequency category, examining aspects such as severity, root causes, and employed fixing patterns.Leveraging the empirical findings, we discuss the broader implications drawn from the study and offer guidance to software developers, facilitating proactive avoidance of such issues in their projects. Hongcheng Fan, Di Liu 0021, Jielun Wu, Yang Feng 0003, Qingkai Shi, Baowen Xu |
Internetware | 2 |
| 2025 | Mining Fine-Grained Code Change Patterns Using Multiple Feature AnalysisabstractMaintaining high code quality is a crucial concern in software development. Existing studies demonstrated that developers frequently face recurrent bugs and adopt similar fix measures, known as code change patterns. As an essential static analysis technique, code pattern mining supports various tasks, including code refactoring, automated program repair, and defect prediction, thus significantly improving software development processes. A prevalent approach to identifying code patterns involves translating code changes to edit actions into a Bag-of-Words (BoW) model. However, when applied to open-source projects, this method exhibits several limitations. For instance, it overlooks function call information and disregards feature word order. This study introduces MIFA, a novel technique for mining code change patterns using multiple feature analysis. MIFA extends existing BoW methods by incorporating analysis of function calls and overall changes in the Abstract Syntax Tree (AST) structure. We selected 20 popular Python projects and evaluated MIFA in both intra-project and cross-project scenarios. The experimental results indicate that: (1) MIFA achieved higher silhouette coefficients and F1 scores compared to other state-of-the-art methods, demonstrating a superior accuracy; (2) MIFA can assist developers in detecting unique change patterns more earlier, with an efficiency improvement of over 40% compared to random sampling. Additionally, we discussed critical parameters for measuring the similarity of code changes, guiding users to apply our method effectively. Di Liu 0021, Yang Feng 0003 |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2024 | Mining Fix Patterns for System Interaction BugsabstractSystem interaction is a fundamental aspect of software development. It involves direct engagement between developers and operating systems, covering tasks such as file management, permission handling, environment dependencies, and parallel development. Accurate system interaction can boost software performance and enhance user experience. On the other hand, improper use often leads to software issues, impacting reliability and stability. Meanwhile, most system interaction bugs typically involve only a minor size of code and follow similar fix patterns. In this paper, we designed a technique to uncover common fix patterns for system interaction bugs. The technique converts bug-fixing behaviors into edit actions, then establishes feature vectors to complete their clustering. Based on this, we present a large-scale study on over 7,800 commits from 37 real Github repositories. We analyzed the results and summarized 19 common fix patterns across 9 categories. Further, we discuss the implications that can support related development, testing, and improvements. These findings will contribute to understanding the essence of system interaction bugs and provide insights for future studies. Di Liu 0021, Yanyan Yan, Hongcheng Fan, Yang Feng 0003 |
Internetware | 1 |
| 2023 | Towards understanding bugs in Python interpreters
Di Liu 0021, Yang Feng 0003, Yanyan Yan, Baowen Xu |
Empir. Softw. Eng. | 1 |
| 2022 | Classifying crowdsourced mobile test reports with image features: An empirical study
Yuying Li 0005, Yang Feng 0003, Di Liu 0021, Chunrong Fang, Zhenyu Chen 0001, Baowen Xu |
J. Syst. Softw. | 4 |
| 2022 | Clustering Crowdsourced Test Reports of Mobile Applications Using Image UnderstandingabstractCrowdsourced testing has been widely used to improve software quality as it can detect various bugs and simulate real usage scenarios. Crowdsourced workers perform tasks on crowdsourcing platforms and present their experiences as test reports, which naturally generates an overwhelming number of test reports. Therefore, inspecting these reports becomes a time-consuming yet inevitable task. In recent years, many text-based prioritization and clustering techniques have been proposed to address this challenge. However, in mobile testing, test reports often consist of only short test descriptions but rich screenshots. Compared with the uncertainty of textual information, well-defined screenshots can often adequately express the mobile application’s activity views. In this paper, by employing image-understanding techniques, we propose an approach for clustering crowdsourced test reports of mobile applications based on both textual and image features to assist the inspection procedure. We employ Spatial Pyramid Matching (SPM) to measure the similarity of the screenshots and use the natural-language-processing techniques to compute the textual distance of test reports. To validate our approach, we conducted an experiment on 6 industrial crowdsourced projects that contain more than 1600 test reports and 1400 screenshots. The results show that our approach is capable of outperforming the baselines by up to 37 percent regarding the APFD metric. Further, we analyze the parameter sensitivity of our approach and discuss the settings for different application scenarios. Di Liu 0021, Yang Feng 0003, James A. Jones, Zhenyu Chen 0001 |
IEEE Trans. Software Eng. | 1 |
| 2018 | Generating descriptions for screenshots to assist crowdsourced testingabstractCrowdsourced software testing has been shown to be capable of detecting many bugs and simulating real usage scenarios. As such, it is popular in mobile-application testing. However in mobile testing, test reports often consist of only some screenshots and short text descriptions. Inspecting and under-standing the overwhelming number of mobile crowdsourced test reports becomes a time-consuming but inevitable task. The paucity and potential inaccuracy of textual information and the well-defined screenshots of activity views within mobile applications motivate us to propose a novel technique to assist developers in understanding crowdsourced test reports by automatically describing the screenshots. To reach this goal, in this paper, we propose a fully automatic technique to generate descriptive words for the well-defined screenshots. We employ the test reports written by professional testers to build up language models. We use the computer-vision technique, namely Spatial Pyramid Matching (SPM), to measure similarities and extract features from the screenshot images. The experimental results, based on more than 1000 test reports from 4 industrial crowdsourced projects, show that our proposed technique is promising for developers to better understand the mobile crowdsourced test reports. Di Liu 0021, Yang Feng 0003, James A. Jones |
SANER | 1 |