Hao Li 0094

dblp:17/5705-94 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0003-4468-5972ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 3 first-author · 8 since 2021
YearPublicationVenuePosition
2026 Mind the Merge: Evaluating the Effects of Token Merging on Pre-Trained Models for Code
Mootez Saad, Hao Li 0094, Tushar Sharma 0001, Ahmed E. Hassan
SANER2
2026 An empirical study of testing practices in open source AI agent frameworks and agentic applications
Mohammed Mehedi Hasan, Hao Li 0094, Emad Fallahzadeh, Gopi Krishnan Rajbahadur, Bram Adams, Ahmed E. Hassan
Empir. Softw. Eng.2
2026 Can we recycle our old models? An empirical evaluation of model selection mechanisms for AIOps solutions
Yingzhe Lyu, Hao Li 0094, Heng Li 0007, Ahmed E. Hassan
Empir. Softw. Eng.2
2026 A systematic literature review of software engineering research on Jupyter notebook
abstract
• This research provides the first comprehensive systematic literature review on software engineering research specifically targeting Jupyter notebooks, identifying 199 primary studies published up to September 2025 and categorizing them into 11 core software engineering topics. • This research reveals that a large portion of the studies have been published outside traditional software engineering venues, with Human-Computer Interaction conferences like ACM Conference on Human Factors in Computing Systems (CHI) being the top publishing venues, highlighting the interdisciplinary nature of Jupyter Notebook research. • This research identifies a reusability gap in existing research, showing that only 82 out of 199 studies offer usable replication packages, and most are hosted on GitHub instead of permanent repositories, which violates open science best practices. • This research identifies that notebook-specific solutions for software engineering issues such as testing, refactoring, and documentation are relatively underexplored. Future directions include resolving duplicated execution numbers, refactoring inter-notebook clones, and generating grouped documentation for coherent-code cells are future directions derived from our study. • This research proposes the integration of modern AI-based solutions into Jupyter notebooks to support various software engineering topics, including code search and code generation. Additionally, future research should leverage advanced AI techniques (e.g., large language models), to improve conversational AI-powered assistants for automated code generation by multi-step workflow automation in data science notebooks. • Although the paper exceeds the recommended length due to the inclusion of detailed tables, figures, and categorized analyses (covering 11 topics and 21 subtopics), we believe that this extended content is essential for clearly and completely reporting our findings. As the first systematic literature review in this domain, we have carefully structured the paper to ensure readability. We believe the length is justified by the value and breadth of this paper’s contributions. Context : Jupyter Notebook has emerged as a versatile tool that transforms how researchers, developers, and data scientists conduct and communicate their work. As the adoption of Jupyter notebooks continues to rise, so does the interest from the software engineering research community in improving the software engineering practices for Jupyter notebooks. Objective : The purpose of this study is to analyze trends, gaps, and methodologies used in software engineering research on Jupyter notebooks. Method : We selected 199 relevant publications up to September 2025, following established systematic literature review guidelines. We explored publication trends, categorized them based on software engineering topics, and reported findings based on those topics. Results : The most popular venues for publishing software engineering research on Jupyter notebooks are related to human-computer interaction instead of traditional software engineering venues. Researchers have addressed a wide range of software engineering topics on notebooks, such as code reuse, readability, and execution environment. Although reusability is one of the research topics for Jupyter notebooks, only 82 of the 199 studies can be reused based on their provided URLs. Additionally, most replication packages are not hosted on permanent repositories for long-term availability and adherence to open science principles. Conclusion : Solutions specific to notebooks for software engineering issues, including testing, refactoring, and documentation, are underexplored. Future research opportunities exist in automatic testing frameworks, refactoring clones between notebooks, and generating group documentation for coherent code cells.
Md. Saeed Siddik, Hao Li 0094, Cor-Paul Bezemer
J. Syst. Softw.2
2026 Towards Refining Developer Questions Using LLM-Based Named Entity Recognition for Developer Chatroom Conversations
abstract
In software engineering chatrooms, communication is often hindered by imprecise questions that cannot be answered. Recognizing key entities (e.g., programming languages and libraries) and user intent (e.g., learning or requesting a review) can be essential for improving question clarity and facilitating better exchange. However, existing research using natural language processing techniques often overlooks these softwarespecific nuances. In this paper, we introduceSoftwarE-specificNamed entity recognition,Intent detection, andResolution classification (SENIR), a labelling approach that leverages a Large Language Model to annotate entities, intents, and resolution status in developer chatroom conversations. To offer quantitative guidance for improving question clarity and resolvability, we build a resolution prediction model that leverages SENIR’s entity and intent labels along with additional predictive features. We evaluate SENIR on the DISCO dataset using a subset of annotated chatroom dialogues. SENIR achieves an 86% F-score for entity recognition, a 71% F-score for intent detection, and an 89% F-score for resolution status classification. Furthermore, our resolution prediction model, tested with various sampling strategies (random undersampling and oversampling with SMOTE) and evaluation methods (5-fold cross-validation, 10-fold cross-validation, and bootstrapping), demonstrates AUC values ranging from 0.7 to 0.8. Key factors influencing resolution include positive sentiment and entities such asProgramming LanguageandUser Variableacross multiple intents, while diagnostic entities (e.g.,Error Name) are more relevant in error-related questions. Moreover, resolution rates vary significantly by intent: questions aboutAPI UsageandAPI Changeachieve higher resolution rates, whereasDiscrepancyandReviewhave lower resolution rates. A Chi-Square analysis confirms the statistical significance of these differences.
Pouya Fathollahzadeh, Mariam El Mezouar, Hao Li 0094, Ying Zou 0001, Ahmed E. Hassan
IEEE Trans. Software Eng.3
2025 Bridging the language gap: an empirical study of bindings for open source machine learning libraries across software package ecosystems
Hao Li 0094, Cor-Paul Bezemer
Empir. Softw. Eng.1
2025 Studying the Impact of TensorFlow and PyTorch Bindings on Machine Learning Software Quality
abstract
Bindings for machine learning frameworks (such as TensorFlow and PyTorch) allow developers to integrate a framework’s functionality using a programming language different from the framework’s default language (usually Python). In this article, we study the impact of using TensorFlow and PyTorch bindings in C#, Rust, Python and JavaScript on the software quality in terms of correctness (training and test accuracy) and time cost (training and inference time) when training and performing inference on five widely used deep learning models. Our experiments show that a model can be trained in one binding and used for inference in another binding for the same framework without losing accuracy. Our study is the first to show that using a non-default binding can help improve machine learning software quality from the time cost perspective compared to the default Python binding while still achieving the same level of correctness.
Hao Li 0094, Gopi Krishnan Rajbahadur, Cor-Paul Bezemer
ACM Trans. Softw. Eng. Methodol.1
2023 An Empirical Study of Yanked Releases in the Rust Package Registry
abstract
Cargo, the software packaging manager of Rust, provides a yank mechanism to support release-level deprecation, which can prevent packages from depending on yanked releases. Most prior studies focused on code-level (i.e., deprecated APIs) and package-level deprecation (i.e., deprecated packages). However, few studies have focused on release-level deprecation. In this study, we investigate how often and how the yank mechanism is used, the rationales behind its usage, and the adoption of yanked releases in the Cargo ecosystem. Our study shows that 9.6% of the packages in Cargo have at least one yanked release, and the proportion of yanked releases kept increasing from 2014 to 2020. Package owners yank releases for other reasons than withdrawing a defective release, such as fixing a release that does not follow semantic versioning or indicating a package is removed or replaced. In addition, we found that 46% of the packages directly adopted at least one yanked release and the yanked releases propagated through the dependency network, which leads to 1.4% of the releases in the ecosystem having unresolved dependencies.
Hao Li 0094, Filipe Roseiro Côgo, Cor-Paul Bezemer
IEEE Trans. Software Eng.1