VLDB 2026 Research / reviewers in the wild / expert
Yanming Yang
dblp:52/2114
· DBLP profile ↗
10ranked-venue papers
8as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Streamlining Java Programming: Uncovering Well-Formed Idioms with IdioMineabstractCode idioms are commonly used patterns, techniques, or practices that aid in solving particular problems or specific tasks across multiple software projects. They can improve code quality, performance, and maintainability, and also promote program standardization and reuse across projects. However, identifying code idioms is significantly challenging, as existing studies have still suffered from three main limitations. First, it is difficult to recognize idioms that span non-contiguous code lines. Second, identifying idioms with intricate data flow and code structures can be challenging. Moreover, they only extract dataset-specific idioms, so common idioms or well-established code/design patterns that are rarely found in datasets cannot be identified. Yanming Yang, Xing Hu 0008, Xin Xia 0001, David Lo 0001, Xiaohu Yang 0001 |
ICSE | 1 |
| 2024 | The Lost World: Characterizing and Detecting Undiscovered Test SmellsabstractTest smell refers to poor programming and design practices in testing and widely spreads throughout software projects. Considering test smells have negative impacts on the comprehension and maintenance of test code and even make code-under-test more defect-prone, it thus has great importance in mining, detecting, and refactoring them. Since Deursen et al. introduced the definition of “test smell”, several studies worked on discovering new test smells from test specifications and software practitioners’ experience. Indeed, many bad testing practices are “observed” by software developers during creating test scripts rather than through academic research and are widely discussed in the software engineering community (e.g., Stack Overflow) [ 70 , 94 ]. However, no prior studies explored new bad testing practices from software practitioners’ discussions, formally defined them as new test smell types, and analyzed their characteristics, which plays a bad role for developers in knowing these bad practices and avoiding using them during test code development. Therefore, we pick up those challenges and act by working on systematic methods to explore new test smell types from one of the most mainstream developers’ Q&A platforms, i.e., Stack Overflow. We further investigate the harmfulness of new test smells and analyze possible solutions for eliminating them. We find that some test smells make it hard for developers to fix failed test cases and trace their failing reasons. To exacerbate matters, we have identified two types of test smells that pose a risk to the accuracy of test cases. Next, we develop a detector to detect test smells from software. The detector is composed of six detection methods for different smell types. These detection methods are both wrapped with a set of syntactic rules based on the code patterns extracted from different test smells and developers’ code styles. We manually construct a test smell dataset from seven popular Java projects and evaluate the effectiveness of our detector on it. The experimental results show that our detector achieves high performance in precision, recall, and F1 score. Then, we utilize our detector to detect smells from 919 real-world Java projects to explore whether the six test smells are prevalent in practice. We observe that these test smells are widely spread in 722 out of 919 Java projects, which demonstrates that they are prevalent in real-world projects. Finally, to validate the usefulness of test smells in practice, we submit 56 issue reports to 53 real-world projects with different smells. Our issue reports achieve 76.4% acceptance by conducting sentiment analysis on developers’ replies. These evaluations confirm the effectiveness of our detector and the prevalence and practicality of new test smell types on real-world projects. Yanming Yang, Xing Hu 0008, Xin Xia 0001, Xiaohu Yang 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2024 | Federated Learning for Software Engineering: A Case Study of Code Clone Detection and Defect PredictionabstractIn various research domains, artificial intelligence (AI) has gained significant prominence, leading to the development of numerous learning-based models in research laboratories, which are evaluated using benchmark datasets. While the models proposed in previous studies may demonstrate satisfactory performance on benchmark datasets, translating academic findings into practical applications for industry practitioners presents challenges. This can entail either the direct adoption of trained academic models into industrial applications, leading to a performance decrease, or retraining models with industrial data, a task often hindered by insufficient data instances or skewed data distributions. Real-world industrial data is typically significantly more intricate than benchmark datasets, frequently exhibiting data-skewing issues, such as label distribution skews and quantity skews. Furthermore, accessing industrial data, particularly source code, can prove challenging for Software Engineering (SE) researchers due to privacy policies. This limitation hinders SE researchers’ ability to gain insights into industry developers’ concerns and subsequently enhance their proposed models. To bridge the divide between academic models and industrial applications, we introduce a federated learning (FL)-based framework calledAlmity. Our aim is to simplify the process of implementing research findings into practical use for both SE researchers and industry developers.Almityenhances model performance on sensitive skewed data distributions while ensuring data privacy and security. It introduces an innovative aggregation strategy that takes into account three key attributes: data scale, data balance, and minority class learnability. This strategy is employed to refine model parameters, thereby enhancing model performance on sensitive skewed datasets. In our evaluation, we employ two well-established SE tasks, i.e., code clone detection and defect prediction, as evaluation tasks. We compare the performance ofAlmityon both machine learning (ML) and deep learning (DL) models against two mainstream training methods, specifically the Centralized Training Method (CTM) and Vanilla Federated Learning (VFL), to validate the effectiveness and generalizability ofAlmity. Our experimental results demonstrate that our framework is not only feasible but also practical in real-world scenarios.Almityconsistently enhances the performance of learning-based models, outperforming baseline training methods across all types of data distributions. Yanming Yang, Xing Hu 0008, Zhipeng Gao 0002, Jinfu Chen 0002, Chao Ni 0001, Xin Xia 0001, David Lo 0001 |
IEEE Trans. Software Eng. | 1 |
| 2023 | C³: Code Clone-Based Identification of Duplicated ComponentsabstractReinventing the wheel is a detrimental programming practice in software development that frequently results in the introduction of duplicated components. This practice not only leads to increased maintenance and labor costs but also poses a higher risk of propagating bugs throughout the system. Despite numerous issues introduced by duplicated components in software, the identification of component-level clones remains a significant challenge that existing studies struggle to effectively tackle. Specifically, existing methods face two primary limitations that are challenging to overcome: 1) Measuring the similarity between different components presents a challenge due to the significant size differences among them; 2) Identifying functional clones is a complex task as determining the primary functionality of components proves to be difficult. Yanming Yang, Ying Zou 0001, Xing Hu 0008, David Lo 0001, Chao Ni 0001, John C. Grundy, Xin Xia 0001 |
ESEC/SIGSOFT FSE | 1 |
| 2022 | MCHPT: A Weakly Supervise Based Merchant Pre-trained Model
Zehua Zeng, Xiaohan She, Xuetao Qiu, Hongfeng Chai, Yanming Yang |
ICONIP (4) | 5 |
| 2022 | Predictive Models in Software Engineering: Challenges and OpportunitiesabstractPredictive models are one of the most important techniques that are widely applied in many areas of software engineering. There have been a large number of primary studies that apply predictive models and that present well-performed studies in various research domains, including software requirements, software design and development, testing and debugging, and software maintenance. This article is a first attempt to systematically organize knowledge in this area by surveying a body of 421 papers on predictive models published between 2009 and 2020. We describe the key models and approaches used, classify the different models, summarize the range of key application areas, and analyze research results. Based on our findings, we also propose a set of current challenges that still need to be addressed in future work and provide a proposed research road map for these opportunities. Yanming Yang, Xin Xia 0001, David Lo 0001, Tingting Bi, John C. Grundy, Xiaohu Yang 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2018 | Structural Function Based Code Clone Detection Using a New Hybrid TechniqueabstractIn this paper, we focus on investigating function based code clone detection and leveraging the structural information to measure the similarity of code fragments in the function level. The method first combines a variant of Abstract Syntax Tree(AST) to achieve more abstract code representations by using defined node types instead of the original node representations, and then adopts a local comparison algorithm, namely Smith Waterman, to calculate the similarity scores of pairs of code fragments in the function level. Experiments conducted over the five open-source datasets show that our method can achieve 92.46% in precision on average, and outperform the comparative algorithms by up to 10.94% and 4.02%, respectively. Meanwhile, experimental results show that our method can achieve 90.73% in precision on average in code clone detection over cross-projects. Yanming Yang, Zhilei Ren, Xin Chen 0032, He Jiang 0001 |
COMPSAC (1) | 1 |
| 2017 | Multi-objective memetic algorithm based on request prediction for dynamic pickup-and-delivery problemsabstractThis paper presents a multi-objective memetic algorithm based on request prediction for route planning in dynamic pickup-and-delivery problems. Historical data are used to predict the occurrence of new dynamic requests, based on which predictive routes are planned and tuned subsequently as the real requests occur. Two objectives namely route length and response time are optimized using multi-objective memetic algorithm that is a synergy of multi-objective genetic algorithm and a locality-sensitive hashing based local search. The proposed algorithm is tested on three benchmark problems and the experimental results demonstrate the efficiency of the algorithm. Yanming Yang, Zexuan Zhu 0001 |
CEC | 1 |
| 2016 | Multi-objective memetic algorithm for solving pickup and delivery problem with dynamic customer requests and traffic informationabstractThis paper formulates one-to-many-to-one pickup and delivery problems with dynamic customer requests and traffic information. A multi-objective memetic algorithm namely prioLSH-MOMA is proposed to solve the problems. The new algorithm is characterized with a priority and locality-sensitive hashing based local search. prioLSH-MOMA is designed to find an optimal route of a dynamic pickup and delivery problem in terms of route length and workload. Particularly, a re-planning strategy is introduced to handle the dynamic information. Priority and locality-sensitive hashing based local search is applied to fine-tune the candidate routes during the evolution process. prioLSH-MOMA is evaluated with two dynamic pickup and delivery problems simulated on real-world maps and the results demonstrate the efficiency of the proposed algorithm. Yanming Yang, Xiaoliang Ma 0001, Zexuan Zhu 0001 |
CEC | 2 |
| 2003 | TSGDB: a database system for tumor suppressor genesabstractUNLABELLED: A Web-based database system was constructed and implemented that contains 174 tumor suppressor genes. The database homepage was created to accommodate these genes in a pull-down window so that each gene can be viewed individually in a separate Web page. Information displayed on each page includes gene name, aliases, source organism, chromosome location, expression cells/tissues, gene structure, protein size, gene functions and major reference sources. Queries to the database can be conducted through a user-friendly interface, and query results are returned in the HTML format on dynamically generated web pages. AVAILABILITY: The database is available at http://www.cise.ufl.edu/~yy1/HTML-TSGDB/Homepage.html (data files also at http://www.patcar.org/Databases/Tumor_Suppressor_Genes) Yanming Yang, Li M. Fu |
Bioinform. | 1 |