VLDB 2026 Research / reviewers in the wild / expert
Yulu Cao
dblp:320/1833
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2024
0000-0002-6623-1165ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 3 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Efficient Construction of Practical Python Call Graphs with Entity Knowledge BaseabstractCall graphs facilitate various tasks in software engineering. However, for the dynamic language Python, the complex language features and external library dependencies pose enormous challenges for building the call graphs of real projects. Some program analysis techniques used for call graph construction in other languages are impractical for Python. In this paper, we present STAR, a practical technique for the construction of Python static call graphs. We reformulate call graph construction as an entity identification task. STAR leverages inter-module summary and cross-project dependencies to construct a fine-grained entity knowledge base to identify the possible nodes and edges of the call graph in the code, and then construct the call graph. Our evaluation of three benchmarks shows that (1) STAR improves recall in three benchmarks compared to three baseline tools. Especially, STAR improves the recall of reachable nodes and reachable edges compared with the state-of-the-art tool by 11.3% and 9.8%, respectively; (2) STAR achieves comparable performance as three baseline tools in execution time and memory usage and is more efficient in large projects; (3) STAR can be effectively used for the task of detecting vulnerability propagation with real-world cases. We expect our results will attract more exploration of practical methods and improve the application of Python call graphs. Yulu Cao, Lin Chen 0015, Zhifei Chen, Jiacheng Zhong, Xiaowei Zhang 0018, Linzhang Wang |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2024 | Multi-Intent Inline Code Comment Generation via Large Language ModelabstractCode comment generation typically refers to the process of generating concise natural language descriptions for a piece of code, which facilitates program comprehension activities. Inline code comments, as a part of code comments, are also crucial for program comprehension. Recently, the emergence of large language models (LLMs) has significantly boosted the performance of natural language processing tasks. This naturally inspires us to explore the performance of the LLMs in the task of inline code comment generation. To this end, we evaluate open-source LLMs on a large-scale dataset and compare the results with the current state-of-the-art methods. Specifically, we explore the model performance in the following scenarios based on the widely used evaluation metrics (i.e. BLEU, Meteor, and ROUGE-L): (1) generation with simple instruction; (2) few-shot-guided generation with random examples selected from the database; (3) few-shot-guided generation with similar examples selected from the database; and (4) adopt the re-ranking strategy for the output of LLMs. Our findings reveal that: (1) under the simple instruction scenario, LLMs could not fully show the potential in the task of inline comment generation compared to the state-of-the-art models; (2) random few-shot leads to a slight improvement; (3) similar few-shot and re-ranking strategy could significantly enhance the performance of LLMs; and (4) for inline comment and code snippet pairs with different intents, why category achieves the best performance and what category achieves relatively poorer performance. That remains consistent across all four scenarios. Our findings shed light on future research directions for using LLMs in inline comment generation tasks. Xiaowei Zhang 0018, Zhifei Chen, Yulu Cao, Lin Chen 0015, Yuming Zhou |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2024 | ICG: A Machine Learning Benchmark Dataset and Baselines for Inline Code Comments Generation TaskabstractAs a fundamental component of software documentation, code comments could help developers comprehend and maintain programs. Several datasets of method header comments have been proposed in previous studies for machine learning-based code comment generation. As part of code comments, inline code comments are also crucial for code understanding activities. However, unlike method header comments written in a standard format and describing the whole method code, inline comments are often written in arbitrary formats by developers due to timelines pressures and different aspects of code snippets in the method are described. Currently, there is no large-scale dataset used for inline comments generation considering these. Hence, this naturally inspires us to explore whether we can construct a dataset to foster machine learning research that not only performs fine-grained noise-cleaning but conducts a taxonomy of inline comments. To this end, we first collect inline comments and code snippets from 8000 Java projects on GitHub. Then, we conduct a manual review to obtain heuristic rules, which could be used to clean the data noise in a fine-grained manner. As a result, we construct a large-scale benchmark dataset named ICG with 5,740,770 pairs of inline comments and code snippets. We then build a comprehensive taxonomy and conduct a statistical and manual analysis to explore the performances of different categories of inline comments, such as helpfulness in code understanding. After that, we provide and compare several baseline models to automatically generate inline comments, such as CodeBERT, to enhance the usability of the benchmark for researchers. The availability of our benchmark and baselines can help develop and validate new inline comment generation methods, which would also further facilitate code understanding activities. Xiaowei Zhang 0018, Lin Chen 0015, Weiqin Zou, Yulu Cao, Hao Ren 0011, Yanhui Li 0001, Yuming Zhou |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2024 | Diagnosis of package installation incompatibility via knowledge base
Yulu Cao, Zhifei Chen, Xiaowei Zhang 0018, Yanhui Li 0001, Lin Chen 0015, Linzhang Wang |
Sci. Comput. Program. | 1 |
| 2024 | Just-in-time identification for cross-project correlated issuesabstractAbstract Issue tracking systems are now prevalent in software development, which would help developers submit and discuss issues to solve development problems on software projects. Most previous studies have been conducted to analyze issue relations within projects, such as recommending similar or duplicate bug issues. However, along with the popularization of co‐developing through multiple projects, many issues are cross‐project correlated (CPC), that is, one issue is associated with another issue in a different project. When developers meet with CPC issues, it may primarily increase the difficulties of solving them because they need information from not only their projects but also other related projects that developers are not familiar with. Identifying a CPC issue as early as possible is a fundamental challenge for both managers and developers to allocate the resources for software maintenance and estimate the effort to solve it. This paper proposes 11 issue metrics of two groups to describe textual summary and reporters' activity, which can be extracted just after the issue was reported. We employ these 11 issue metrics to construct just‐in‐time (JIT) prediction models to identify CPC issues. To evaluate the effect of CPC issue prediction models, we conduct experiments on 16 open‐source data science and deep learning projects and compare our prediction model with two baseline models based on textual features (i.e., Term Frequency‐Inverse Document Frequency [TF‐IDF] and Word Embedding), which are commonly adopted by previous studies on issue prediction. The results show that the JIT prediction model based on issue metrics has significantly improved the performance of CPC issue prediction under two evaluation indicators, Matthew's correlation coefficient (MCC) and F1. In addition, we find that the prediction model is more suitable for large‐scale complex core projects in the open‐source ecosystem. Hao Ren 0011, Yanhui Li 0001, Lin Chen 0015, Yulu Cao, Xiaowei Zhang 0018, Changhai Nie |
J. Softw. Evol. Process. | 4 |
| 2023 | Effective Recommendation of Cross-Project Correlated Issues based on Issue MetricsabstractThe calling relationship between projects becomes complicated as the number of open-source projects increases. Different issues across projects can also be related, referred to as cross-project correlated issues (CPCIs), and bring new challenges for developers to fix these issues. When solving these CPCIs, developers have to accurately locate the source code that causes it in the current project and also needs to know the related issues in other projects. However, few studies have proposed specific methods to help developers effectively address these CPCIs, i.e., find related issues for CPCIs. Hao Ren 0011, Mingliang Ma, Xiaowei Zhang 0018, Yulu Cao, Changhai Nie |
Internetware | 4 |
| 2023 | Towards Better Dependency Scope Settings in Maven ProjectsabstractThe emergence of build automation tools with dependency management features has significantly impacted software development. However, in the configuration process, improper settings of some configuration items, such as the dependency scope setting, may cause severe problems in the development process. Improper setting of dependency scope may cause problems such as missing dependencies and redundant dependencies, and may even spread the problem to the downstream of the software ecosystem. Lin Chen 0015, Yulu Cao, Yanhui Li 0001, Yuming Zhou |
Internetware | 3 |
| 2023 | Towards Better Dependency Management: A First Look at Dependency Smells in Python ProjectsabstractManaging cross-project dependencies is tricky in modern software development. A primary way to manage dependencies is using dependency configuration files, which brings convenience to the entire software ecosystem, including developers, maintainers, and users. However, developers may introduce dependency smells if dependency configuration files are not well written and maintained. Dependency smells are recurring violations of dependency management in dependency configuration files and can potentially lead to severe consequences. This paper provides an in-depth look at three dependency smells, namely,Missing Dependency,Bloated Dependency, andVersion Constraint Inconsistencyin Python projects. First, we implement a tool calledPythonCross-projectDependency- PyCD to accurately extract dependency information from configuration files. The evaluation result on 212 Python projects shows that PyCD outperforms state-of-the-art tools. Then, we make an empirical study for three dependency smells in 132 Python projects to investigate the pervasiveness, causes, and evolution. The results show that: 1) dependency smells are prevalent in Python projects and exist inconsistently in different projects; 2) dependency smells are introduced into Python projects for different reasons, mainly due to the problems of synchronous update and collaborative development; and 3) dependency smells can be removed with different patterns according to different dependency smells. Furthermore, we report and get responses for 40 harmful dependency smell instances, 34 of which have been responded that these dependency smells do exist in the projects, and 10 instances are fixed or under process. The feedback from developers indicates that dependency smells can have a negative impact on project maintenance. Our study highlights that these dependency smells deserve the attention of developers. Yulu Cao, Lin Chen 0015, Wanwangying Ma, Yanhui Li 0001, Yuming Zhou, Linzhang Wang |
IEEE Trans. Software Eng. | 1 |