VLDB 2026 Research / reviewers in the wild / expert
Haowei Quan
dblp:314/6781
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2026
0000-0003-0863-7973ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Empirical Study of Vulnerabilities in Python Packages and Their DetectionabstractContextIn the rapidly evolving software development landscape, Python stands out for its simplicity, versatility, and extensive ecosystem. Python packages, as units of organization, reusability, and distribution, have become a pressing concern, highlighted by the considerable number of vulnerability reports. As a scripting language, Python often cooperates with other programming languages for performance or interoperability. This also adds complexity to the vulnerabilities inherent to Python packages, and the effectiveness of current vulnerability detection tools remains underexplored in the research community.ObjectivesTo bridge this gap, we present PyVul, the first comprehensive benchmark suite of Pythonpackage vulnerabilities. We use this benchmark to conduct an empirical study that characterizes these vulnerabilities and evaluates the limitations of state-of-the-art detection tools.MethodsWe collect real-world vulnerability reports from GitHub Advisories, Snyk, and Huntr, and curate our benchmark at both the commit level and function level. To improve accuracy, we propose LLM-VDC, a large language model–assisted cleansing method. Based on PyVul, we systematically analyze vulnerabilities and assess the capabilities of both rule-based and machine learning–based detectors.ResultsAfter cleansing, PyVul achieves an accuracy of 100% at the commit level with 1,157 repository snapshots, and 94.0% at the function level with 2,082 vulnerable functions, establishing it as the most precise automatically collected Python vulnerability benchmark. Our empirical analysis reveals that current rule-based vulnerability detectors suffer from mismatches between their assumptions and real-world security scenarios, and limited support for high-order vulnerabilities, cross-language interactions, and Python’s unique language features. On the other hand, ML-based detectors suffer from their inability to reach the necessary context.ConclusionA significant discrepancy exists between the capabilities of existing tools and the demands of effectively identifying real-world security issues in Python packages. PyVul provides a solid foundation for advancing vulnerability research and tool development in this domain. Haowei Quan, Junjie Wang 0007, Terry Yue Zhuo, Xiao Chen 0002, Xiaoning Du 0001 |
MSR | 1 |
| 2024 | Neural Library Recommendation by Embedding Project-Library Knowledge GraphabstractThe prosperity of software applications brings fierce market competition to developers. Employing third-party libraries (TPLs) to add new features to projects under development and to reduce the time to market has become a popular way in the community. However, given the tremendous TPLs ready for use, it is challenging for developers to effectively and efficiently identify the most suitable TPLs. To tackle this obstacle, we propose an innovative approach named PyRec to recommend potentially useful TPLs to developers for their projects. Taking Python project development as a use case, PyRec embeds Python projects, TPLs, contextual information, and relations between those entities into a knowledge graph. Then, it employs a graph neural network to capture useful information from the graph to make TPL recommendations. Different from existing approaches, PyRec can make full use of not only project-library interaction information but also contextual information to make more accurate TPL recommendations. Comprehensive evaluations are conducted based on 12,421 Python projects involving 963 TPLs, 9,675 extra entities, 121,474 library usage records, and 73,277 contextual records. Compared with five representative approaches, PyRec improves the recommendation performance significantly in all cases. Bo Li 0103, Haowei Quan, Jiawei Wang 0003, Haipeng Cai, Yuan Miao 0001, Yun Yang 0001, Li Li 0029 |
IEEE Trans. Software Eng. | 2 |