EDBT 2026 Demo / reviewers in the wild / expert
Di Wu 0014
dblp:52/328-14
· DBLP profile ↗
25ranked-venue papers
14as first author
18since 2021 · last 2025
0000-0003-1096-7074ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 22 · 14 first-author · 16 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Using Prompt Tuning to Identify Relevant API Knowledge from API Tutorial and Stack OverflowabstractAPI tutorials and Stack Overflow (SO) are crucial API learning resources. API tutorials help developers understand API usage in general contexts, while SO explains API usage in specific programming tasks. Using both API tutorials and SO provides more API knowledge. We treat a tutorial fragment or a SO Question and Answering (Q&A for short) pair as a knowledge item (KI for short). Discovering relevant KIs of the APIs helps developers understand and learn how to use APIs. However, existing relevant KIs identification approaches mainly focus on either API tutorials or SO. Furthermore, these approaches do not take into account code in KIs. In this paper, we propose PTIRK, a novel approach using prompt tuning to identify relevant KIs of API from both API tutorials and SO. Firstly, we extract textual descriptions and code snippets from KIs and match them with the API class name. Then, prompt tuning is used to tune BERT for text in KIs and CodeBERT for code in KIs. After that, we combine tuned models by using a gating fusion strategy. Finally, the relevant KIs of API can be identified by the trained model. We evaluate PTIRK on Java and Android datasets with 10,072 samples. Experimental results show that PTIRK outperforms state-of-the-art approaches on both datasets, and the user study further confirms its practical effectiveness. Di Wu 0014, Yanfei Sun |
QRS | 1 |
| 2025 | Multi-view learning based on product and process metrics for software defect prediction
Ying Sun 0023, Fei Wu 0004, Di Wu 0014, Xiaoyuan Jing, Yanfei Sun |
Appl. Intell. | 3 |
| 2025 | MITU: Locating relevant tutorial fragments of APIs with multi-source API knowledge
Di Wu 0014, Hongyu Zhang 0002, Yang Feng 0003, Zhenjiang Dong |
J. Syst. Softw. | 1 |
| 2025 | PFBL: Prototype-Based Fully Balanced Learning for Multimodal Fake News DetectionabstractMultimodal fake news detection (MFND) has received widespread attention. As a typical multimodal task, MFND is troubled by modality imbalance, where the dominant modality suppresses other modalities during the optimization process. Meanwhile, the social credibility maintains the amount of fake news less than that of real news, which causes class imbalance. We call MFND with intertwined influence of modality imbalance and class imbalance as dual imbalanced MFND, i.e., DI-MFND, which has not been well studied. In this article, we propose an approach called prototype-based fully balanced learning (PFBL) for DI-MFND. Specifically, our model contains three main parts. 1) The prototype-based modality-balanced learning (PMB) part, which constructs modality prototypes for each modality. It designs the prototype-based unimodal loss to enhance the intraclass compactness of the poorly performing modality and a constrained gradient optimization strategy to suppress the dominant modality for optimization. 2) The prototype-based class-balanced learning (PCB) part constructs class prototypes. A prototype-oriented discriminative enhancement loss is designed to effectively align samples with corresponding prototypes and increase the inter-class distance, thus enhancing the separability of different categories. 3) The multimodal fusion and classification (MFC) part employs a cross-transformer block to perform interaction between modalities and further integrates features of different modalities by attention mechanism. We propose a joint updating strategy for modality prototypes and class prototypes. Extensive experiments on three widely-used news datasets demonstrate that our approach outperforms state-of-the-art approaches. Fei Wu 0004, Zhe-Ying Deng, Di Wu 0014, Chao Lan, Yimu Ji 0001, Xiaoyuan Jing |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Automatic recognizing relevant fragments of APIs using API references
Di Wu 0014, Yang Feng 0003, Hongyu Zhang 0002, Baowen Xu |
Autom. Softw. Eng. | 1 |
| 2024 | The future of API analytics
Di Wu 0014, Hongyu Zhang 0002, Yang Feng 0003, Zhenjiang Dong, Ying Sun 0023 |
Autom. Softw. Eng. | 1 |
| 2023 | Leveraging Stack Overflow to detect relevant tutorial fragments of APIs
Di Wu 0014, Xiaoyuan Jing, Hongyu Zhang 0002, Yuming Zhou, Baowen Xu |
Empir. Softw. Eng. | 1 |
| 2023 | Retrieving API Knowledge from Tutorials and Stack Overflow Based on Natural Language QueriesabstractWhen encountering unfamiliar APIs, developers tend to seek help from API tutorials and Stack Overflow (SO). API tutorials help developers understand the API knowledge in a general context, while SO often explains the API knowledge in a specific programming task. Thus, tutorials and SO posts together can provide more API knowledge. However, it is non-trivial to retrieve API knowledge from both API tutorials and SO posts based on natural language queries. Two major problems are irrelevant API knowledge in two different resources and the lexical gap between the queries and documents. In this article, we regard a fragment in tutorials and a Question and Answering (Q&A) pair in SO as a knowledge item (KI). We generate ⟨ API, FRA ⟩ pairs (FRA stands for fragment) from tutorial fragments and APIs and build ⟨ API, QA ⟩ pairs based on heuristic rules of SO posts. We fuse ⟨ API, FRA ⟩ pairs and ⟨ API, QA ⟩ pairs to generate API knowledge (AK for short) datasets, where each data item is an ⟨ API, KI ⟩ pair. We propose a novel approach, called PLAN, to automatically retrieve API knowledge from both API tutorials and SO posts based on natural language queries. PLAN contains three main stages: (1) API knowledge modeling, (2) query mapping, and (3) API knowledge retrieving. It first utilizes a deep-transfer-metric-learning-based relevance identification (DTML) model to effectively find relevant ⟨ API, KI ⟩ pairs containing two different knowledge items (⟨ API, QA ⟩ pairs and ⟨ API, FRA ⟩ pairs) simultaneously. Then, PLAN generates several potential APIs as a way to reduce the lexical gap between the query and ⟨ API, KI ⟩ pairs. According to potential APIs, we can select relevant ⟨ API, KI ⟩ pairs to generate potential results. Finally, PLAN returns a list of ranked ⟨ API, KI ⟩ pairs that are related to the query. We evaluate the effectiveness of PLAN with 270 queries on Java and Android AK datasets containing 10,072 ⟨ API, KI ⟩ pairs. Our experimental results show that PLAN is effective and outperforms the state-of-the-art approaches. Our user study further confirms the effectiveness of PLAN in locating useful API knowledge. Di Wu 0014, Xiaoyuan Jing, Hongyu Zhang 0002, Yang Feng 0003, Yuming Zhou, Baowen Xu |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2022 | Training Data Debugging for the Fairness of Machine Learning SoftwareabstractWith the widespread application of machine learning (ML) software, especially in high-risk tasks, the concern about their unfairness has been raised towards both developers and users of ML software. The unfairness of ML software indicates the software behavior affected by the sensitive features (e.g., sex), which leads to biased and illegal decisions and has become a worthy problem for the whole software engineering community. Yanhui Li 0001, Linghan Meng, Lin Chen 0015, Li Yu 0008, Di Wu 0014, Yuming Zhou, Baowen Xu |
ICSE | 5 |
| 2022 | An Empirical Study on the Impact of Python Dynamic Typing on the Project MaintenanceabstractPython is a popular typical dynamic programming language. In Python, dynamic typing is one of the most critical dynamic features. The lack of type information is likely to hinder the maintenance of Python projects. However, existing work has seldom focused on studying the impact of Python dynamic typing on project maintenance. This paper focuses on the two most common practices of Python dynamic typing, i.e. inconsistent-type assignments (ITA) and inconsistent variable types (IVT). Two approaches are proposed to identify ITA and IVT, i.e. identifying ITA by analyzing Abstract Syntax Trees and comparing identifiers types and identifying IVT by constructing a type dependency graph. In empirical experiments, we first locate the usage of ITA and IVT in 10 open-source Python projects. Then, we investigate the relations between the occurrence of ITA and IVT and the results of maintenance tasks. The study results show that projects are more prone to change as the number of dynamic typing identifiers increases. There is a weak connection between change-proneness and variable dynamic typing. There is a high probability that maintenance time and the acceptance of commits decrease as dynamic typing identifiers increase in projects. These results implicate that dynamic and static variables should be divided while developing new programming languages. Dynamic typing identifiers may not be the direct root causes for most software bugs. The categories of these bugs are worth exploring. Xinmeng Xia, Yanyan Yan, Xincheng He, Di Wu 0014, Lei Xu 0003, Baowen Xu |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2022 | How higher order mutant testing performs for deep learning models: A fine-grained evaluation of test effectiveness and efficiency improved from second-order mutant-classification tuples
Yanhui Li 0001, Weijun Shen, Tengchao Wu, Lin Chen 0015, Di Wu 0014, Yuming Zhou, Baowen Xu |
Inf. Softw. Technol. | 5 |
| 2021 | Measuring Discrimination to Boost Comparative Testing for Multiple Deep Learning ModelsabstractThe boom of DL technology leads to massive DL models built and shared, which facilitates the acquisition and reuse of DL models. For a given task, we encounter multiple DL models available with the same functionality, which are considered as candidates to achieve this task. Testers are expected to compare multiple DL models and select the more suitable ones w.r.t. the whole testing context. Due to the limitation of labeling effort, testers aim to select an efficient subset of samples to make an as precise rank estimation as possible for these models. To tackle this problem, we propose Sample Discrimination based Selection (SDS) to select efficient samples that could discriminate multiple models, i.e., the prediction behaviors (right/wrong) of these samples would be helpful to indicate the trend of model performance. To evaluate SDS, we conduct an extensive empirical study with three widely-used image datasets and 80 real world DL models. The experiment results show that, compared with state-of-the-art baseline methods, SDS is an effective and efficient sample selection method to rank multiple DL models. Linghan Meng, Yanhui Li 0001, Lin Chen 0015, Di Wu 0014, Yuming Zhou, Baowen Xu |
ICSE | 5 |
| 2021 | Leveraging Stack Overflow to Detect Relevant Tutorial Fragments of APIsabstractDevelopers often use learning resources such as API tutorials and Stack Overflow (SO) to learn how to use an unfamiliar API. An API tutorial can be divided into a number of consecutive units that describe the same topic, denoted as tutorial fragments. We consider a tutorial fragment explaining the API usage knowledge as a relevant fragment of the API. Discovering relevant tutorial fragments of APIs can facilitate API understanding and learning. However, existing approaches, based on supervised or unsupervised approaches, often suffer from either high manual efforts or lack of consideration of the relevance information. In this paper, we propose a novel approach, called SO2RT, to detect relevant tutorial fragments of APIs based on SO posts. SO2RT first automatically extracts relevant and irrelevant 〈API,QA〉 pairs based on heuristic rules of SO, and constructs 〈API, FRA〉 pairs (FRA stands out fragment) by using tutorial fragments and APIs. SO2RT then trains a semi-supervised transfer learning based detection model, which can transfer the API usage knowledge in SO Q&A pairs to tutorial fragments by utilizing the easy-to-extract relevance of 〈API, QA〉 pairs. Finally, relevant fragments of APIs can be discovered by consulting the trained model. In this way, the effort for labeling the relevance between tutorial fragments and APIs can be reduced. We evaluate SO2RT on Java and Android datasets containing 21,008 〈API, QA〉 pairs. Experimental results show that SO2RT improves the state-of-the-art approaches in terms of F-Measure on both datasets. Our user study further confirms the effectiveness of SO2RT in practice. Di Wu 0014, Xiaoyuan Jing, Hongyu Zhang 0002, Yuming Zhou, Baowen Xu |
SANER | 1 |
| 2021 | Generating API tags for tutorial fragments from Stack Overflow
Di Wu 0014, Xiaoyuan Jing, Hongyu Zhang 0002, Baowen Xu |
Empir. Softw. Eng. | 1 |
| 2021 | Recommending Relevant Tutorial Fragments for API-Related Natural Language QuestionsabstractApplication Programming Interface (API) tutorial is an important API learning resource. To help developers learn APIs, an API tutorial is often split into a number of consecutive units that describe the same topic (i.e. tutorial fragment). We regard a tutorial fragment explaining an API as a relevant fragment of the API. Automatically recommending relevant tutorial fragments can help developers learn how to use an API. However, existing approaches often employ supervised or unsupervised manner to recommend relevant fragments, which suffers from much manual annotation effort or inaccurate recommended results. Furthermore, these approaches only support developers to input exact API names. In practice, developers often do not know which APIs to use so that they are more likely to use natural language to describe API-related questions. In this paper, we propose a novel approach, called Tutorial Fragment Recommendation (TuFraRec), to effectively recommend relevant tutorial fragments for API-related natural language questions, without much manual annotation effort. For an API tutorial, we split it into fragments and extract APIs from each fragment to build API-fragment pairs. Given a question, TuFraRec first generates several clarification APIs that are related to the question. We use clarification APIs and API-fragment pairs to construct candidate API-fragment pairs. Then, we design a semi-supervised metric learning (SML)-based model to find relevant API-fragment pairs from the candidate list, which can work well with a few labeled API-fragment pairs and a large number of unlabeled API-fragment pairs. In this way, the manual effort for labeling the relevance of API-fragment pairs can be reduced. Finally, we sort and recommend relevant API-fragment pairs based on the recommended strategy. We evaluate TuFraRec on 200 API-related natural language questions and two public tutorial datasets (Java and Android). The results demonstrate that on average TuFraRec improves NDCG@5 by 0.06 and 0.09, and improves Mean Reciprocal Rank (MRR) by 0.07 and 0.09 on two tutorial datasets as compared with the state-of-the-art approach. Di Wu 0014, Xiaoyuan Jing, Xiaohui Kong, Jifeng Xuan |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2021 | Boundary sampling to boost mutation testing for deep learning models
Weijun Shen, Yanhui Li 0001, Yuanlei Han, Lin Chen 0015, Di Wu 0014, Yuming Zhou, Baowen Xu |
Inf. Softw. Technol. | 5 |
| 2021 | Similarity-Maintaining Privacy Preservation and Location-Aware Low-Rank Matrix Factorization for QoS Prediction Based Web Service RecommendationabstractWeb service recommendation plays an important role in building service-oriented systems. QoS-based Web service recommendation has recently gained much attention for providing a promising way to help users find high-quality services. To accurately predict the QoS values of candidate Web services, Web service recommendation systems usually need to collect historical QoS data from users, which will potentially pose a threat to the user's privacy. However, how to simultaneously protect user's privacy and make an accurate prediction has not been well studied. By taking these two aspects into consideration, we propose a novel QoS prediction approach for Web service recommendation in this paper. Specifically, we first design a similarity-maintaining privacy preservation (SPP) strategy, which aims to protect the user's privacy and maintain the utility of user data in the meanwhile. Then, we propose a location-aware low-rank matrix factorization (LLMF) algorithm, which employs the L1L1-norm low-rank matrix factorization to improve the model's robustness, and combines the matrix factorization model with two kinds of location information (continent, longitude and latitude) in the prediction process. Experimental results on two publicly available real-world Web service QoS datasets demonstrate the effectiveness of our privacy-preserving QoS prediction approach. Xiaoke Zhu, Xiaoyuan Jing, Di Wu 0014, Zhenyu He 0001, Jicheng Cao, Dong Yue 0001, Lina Wang 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2021 | An Empirical Study on Heterogeneous Defect Prediction ApproachesabstractSoftware defect prediction has always been a hot research topic in the field of software engineering owing to its capability of allocating limited resources reasonably. Compared with cross-project defect prediction (CPDP), heterogeneous defect prediction (HDP) further relaxes the limitation of defect data used for prediction, permitting different metric sets to be contained in the source and target projects. However, there is still a lack of a holistic understanding of existing HDP studies due to different evaluation strategies and experimental settings. In this paper, we provide an empirical study on HDP approaches. We review the research status systematically and compare the HDP approaches proposed from 2014 to June 2018. Furthermore, we also investigate the feasibility of HDP approaches in CPDP. Through extensive experiments on 30 projects from five datasets, we have the following findings: (1) metric transformation-based HDP approaches usually result in better prediction effects, while metric selection-based approaches have better interpretability. Overall, the HDP approach proposed by Liet al.(CTKCCA) currently has the best performance. (2) Handling class imbalance problems can boost the prediction effects, but the improvements are usually limited. In addition, utilizing mixed project data cannot improve the performance of HDP approaches consistently since the label information in the target project is not used effectively. (3) HDP approaches are feasible for cross-project defect prediction in which the source and target projects have the same metric set. Xiaoyuan Jing, Zhiqiang Li 0003, Di Wu 0014, Zhiguo Huang |
IEEE Trans. Software Eng. | 4 |
| 2020 | How C++ Templates Are Used for Generic Programming: An Empirical Study on 50 Open Source SystemsabstractGeneric programming is a key paradigm for developing reusable software components. The inherent support for generic constructs is therefore important in programming languages. As for C++, the generic construct, templates, has been supported since the language was first released. However, little is currently known about how C++ templates are actually used in developing real software. In this study, we conduct an experiment to investigate the use of templates in practice. We analyze 1,267 historical revisions of 50 open source systems, consisting of 566 million lines of C++ code, to collect the data of the practical use of templates. We perform statistical analyses on the collected data and produce many interesting results. We uncover the following important findings: (1) templates are practically used to prevent code duplication, but this benefit is largely confined to a few highly used templates; (2) function templates do not effectively replace C-style generics, and developers with a C background do not show significant preference between the two language constructs; (3) developers seldom convert dynamic polymorphism to static polymorphism by using CRTP (Curiously Recursive Template Pattern); (4) the use of templates follows a power-law distribution in most cases, and C++ developers who prefer using templates are those without other language background; (5) C developer background seems to override C++ project guidelines. These findings are helpful not only for researchers to understand the tendency of template use but also for tool builders to implement better tools to support generic programming. Lin Chen 0015, Di Wu 0014, Wanwangying Ma, Yuming Zhou, Baowen Xu, Hareton K. N. Leung |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2017 | Multi-Kernel Low-Rank Dictionary Pair Learning for Multiple Features Based Image ClassificationabstractDictionary learning (DL) is an effective feature learning technique, and has led to interesting results in many classification tasks. Recently, by combining DL with multiple kernel learning (which is a crucial and effective technique for combining different feature representation information), a few multi-kernel DL methods have been presented to solve the multiple feature representations based classification problem. However, how to improve the representation capability and discriminability of multi-kernel dictionary has not been well studied. In this paper, we propose a novel multi-kernel DL approach, named multi-kernel low-rank dictionary pair learning (MKLDPL). Specifically, MKLDPL jointly learns a kernel synthesis dictionary and a kernel analysis dictionary by exploiting the class label information. The learned synthesis and analysis dictionaries work together to implement the coding and reconstruction of samples in the kernel space. To enhance the discriminability of the learned multi-kernel dictionaries, MKLDPL imposes the low-rank regularization on the analysis dictionary, which can make samples from the same class have similar representations. We apply MKLDPL for multiple features based image classification task. Experimental results demonstrate the effectiveness of the proposed approach. Xiaoke Zhu, Xiaoyuan Jing, Fei Wu 0004, Di Wu 0014, Li Cheng 0006, Ruimin Hu |
AAAI | 4 |
| 2016 | An extensive empirical study on C++ concurrency constructs
Di Wu 0014, Lin Chen 0015, Yuming Zhou, Baowen Xu |
Inf. Softw. Technol. | 1 |
| 2015 | An Empirical Study on C++ Concurrency ConstructsabstractNowadays concurrent programming is in large demand. The inherent support for concurrency is therefore increasingly important in programming languages. As for C++, an abundance of standard concurrency constructs have been supported since C++11. However, to date there is little work investigating how these constructs are actually used in developing real software. In this paper, we perform an empirical study to investigate the adoption of C++ concurrency constructs in open-source applications, with the goal to provide insightful information for practitioners to use concurrency constructs efficiently. To this end, we analyze 127 open-source applications that adopt C++ concurrency constructs, comprising 34 million lines of C++ code, to conduct the experiment. The experimental results show that: (1) to implement concurrency code, thread-based constructs are significantly more often used than atomics-based constructs and task-based constructs; (2) to manage synchronization, lock-based constructs are significantly more often used than lock-free constructs and blocking constructs; (3) among the key thread-based constructs and task-based constructs (i.e. mutex, promise, and future), there is not a construct significantly more commonly misused than others; (4) small-size applications introduce concurrency constructs more intensively and more quickly than medium-size applications and large-size applications; and (5) an increasing use of standard concurrency constructs does not result in a substantially decreasing use of unstandardized concurrency constructs. Based on these findings, we make actionable suggestions for language designers, developers, and novices to assist them in designing and using C++ concurrency constructs. Di Wu 0014, Lin Chen 0015, Yuming Zhou, Baowen Xu |
ESEM | 1 |
| 2015 | How do developers use C++ libraries? An empirical studyabstractC++ libraries provide an abundance of reusable components for writing high-quality programs and are thus widely adopted by software developers.However, to date there is little work investigating how these libraries are actually used in real software.In this paper, we perform an empirical study to investigate the adoption of C++ standard libraries in open-source applications, with the goal to provide actionable information for developers to help them employ libraries more efficiently.To this end, we analyze 379 historical revisions of 30 applications, containing 149 million lines of C++ code, to conduct the experiment.The experimental results show that: (1) three standard libraries (i.e.Containers Library, Utilities Library, and Strings Library) are significantly more often used than other libraries; (2) the new libraries of C++11 (i.e.Regular Expressions Library, Atomic Operations Library, and Thread Support Library) are significantly less often used than the formerlyestablished libraries; (3) the deprecated library constructs (i.e. auto pointers, function objects, and array I/O operations) are not used at a declining frequency; and (4) applications with a larger size do not adopt libraries more frequently.Based on these results, we propose four suggestions, which could help developers learn and use C++ libraries in an efficient way. Di Wu 0014, Lin Chen 0015, Yuming Zhou, Baowen Xu |
SEKE | 1 |
| 2015 | A metrics-based comparative study on object-oriented programming languagesabstractThere has been a long debate on which programming language can help write better object-oriented programs.However, to date little response is given to this issue with empirical evidence.In this paper, we perform a comparative study on C++, C#, and Java programs by using object-oriented metrics, which comprise measures for class size, complexity, coupling, cohesion, inheritance, encapsulation, polymorphism, and reusability.Our experiment is conducted on 78 tasks in Rosetta Code, a code repository providing solutions to the same programming tasks in different languages.The experimental results show that: (1) C++ classes are significantly larger than C# and Java classes in size, but their complexity does not differ significantly; (2) C# classes are significantly more likely to be coupled than C++ and Java classes through inter-class method invocations instead of direct data access; (3) C# and Java classes tend to be more cohesive than C++ classes; (4) C# and Java significantly outperform C++ in building deep inheritance trees; and (5) programs written in C++, C#, and Java do not show a significant difference in class encapsulation, polymorphism, and reusability.These findings could help practitioners choose suitable languages to develop object-oriented systems. 1 Di Wu 0014, Lin Chen 0015, Yuming Zhou, Baowen Xu |
SEKE | 1 |
| 2014 | An empirical study on the adoption of C++ templates: Library templates versus user defined templates
Di Wu 0014, Lin Chen 0015, Yuming Zhou, Baowen Xu |
SEKE | 1 |