EDBT 2026 Demo / reviewers in the wild / expert
Wenhua Yang 0001
dblp:21/11181-1
· DBLP profile ↗
39ranked-venue papers
11as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 34 · 7 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 2 since 2021Computer networks · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Defending neural code understanding models by eliminating backdoors
Yu Zhou 0010, Guang Yang 0019, Xiangyu Zhang 0005, Wenhua Yang 0001, Taolue Chen 0001 |
Autom. Softw. Eng. | 5 |
| 2026 | Detecting duplicate vulnerability records across databases
Kangliang Zhu, Wenhua Yang 0001, Minxue Pan, Yu Zhou 0010 |
Sci. Comput. Program. | 2 |
| 2025 | Understanding Feature Request Practice on GitHub via a Large-Scale Empirical StudyabstractFeature requests are a key communication mechanism on GitHub, enabling users and developers to collaboratively shape the direction of open-source projects. Feature requests are prevalent and important, but have been underexplored in existing studies. There is limited understanding of how they are labeled, how they evolve, and how they are resolved. A deeper understanding of feature requests is critical, not only for improving issue triage and project management but also for fostering more effective collaboration within open-source communities. In this work, we present the first systematic and large-scale empirical study of feature requests. Drawing on 1.4 million issues from 825 GitHub repositories, we examine how feature requests are labeled, how their submission and backlog patterns change over a project’s lifecycle, how they differ from other types of issues in terms of resolution and engagement, and what factors contribute to their successful handling. Our findings reveal that labeling practices are often inconsistent across projects, that feature requests follow distinct temporal trends, and that those which are lengthy and contain large code snippets tend to be more difficult to resolve. By contrast, concise and clearly defined requests, particularly those submitted by experienced contributors and accompanied by active discussions, are more likely to be addressed. This study underscores the challenges of managing feature requests at scale and provides practical insights for maintainers, contributors, and researchers. To support future work in this area, we publicly release our dataset and analysis results. Wenhua Yang 0001, Minxue Pan, Yu Zhou 0010 |
ASE | 2 |
| 2025 | Detecting data manipulation errors in android applications using scene-guided exploration
Yu Zhou 0010, Wenhua Yang 0001, Taolue Chen 0001, Harald C. Gall |
Empir. Softw. Eng. | 3 |
| 2024 | A Two-Stage Approach for GitHub Issue Links Identification and ClassificationabstractEffective issue management is critical for the success of open-source projects on GitHub. However, the platform currently lacks the capability to identify implicit links between issues, complicating the management process. In this study, we propose a machine learning-based approach to identify and classify these links. Preliminary experimental results demonstrate the effectiveness of our approach in identifying issue links, indicating its potential to enhance issue tracking in large-scale GitHub projects. Yingying He, Wenhua Yang 0001 |
APSEC | 2 |
| 2024 | Comprehensive Semantic Repair of Obsolete GUI Test Scripts for Mobile ApplicationsabstractGraphical User Interface (GUI) testing is one of the primary approaches for testing mobile apps. Test scripts serve as the main carrier of GUI testing, yet they are prone to obsolescence when the GUIs change with the apps' evolution. Existing repair approaches based on GUI layouts or images prove effective when the GUI changes between the base and updated versions are minor, however, they may struggle with substantial changes. In this paper, a novel approach named COSER is introduced as a solution to repairing broken scripts, which is capable of addressing larger GUI changes compared to existing methods. COSER incorporates both external semantic information from the GUI elements and internal semantic information from the source code to provide a unique and comprehensive solution. The efficacy of COSER was demonstrated through experiments conducted on 20 Android apps, resulting in superior performance when compared to the state-of-the-art tools METER and GUIDER. In addition, a tool that implements the COSER approach is available for practical use and future research. Shaoheng Cao, Minxue Pan, Yu Pei 0001, Wenhua Yang 0001, Tian Zhang 0001, Linzhang Wang, Xuandong Li |
ICSE | 4 |
| 2024 | Deeply Reinforcing Android GUI Testing with Deep Reinforcement LearningabstractAs the scale and complexity of Android applications continue to grow in response to increasing market and user demands, quality assurance challenges become more significant. While previous studies have demonstrated the superiority of Reinforcement Learning (RL) in Android GUI testing, its effectiveness remains limited, particularly in large, complex apps. This limitation arises from the ineffectiveness of Tabular RL in learning the knowledge within the large state-action space of the App Under Test (AUT) and from the suboptimal utilization of the acquired knowledge when employing more advanced RL techniques. To address such limitations, this paper presents DQT, a novel automated Android GUI testing approach based on deep reinforcement learning. DQT preserves widgets' structural and semantic information with graph embedding techniques, building a robust foundation for identifying similar states or actions and distinguishing different ones. Moreover, a specially designed Deep Q-Network (DQN) effectively guides curiosity-driven exploration by learning testing knowledge from runtime interactions with the AUT and sharing it across states or actions. Experiments conducted on 30 diverse open-source apps demonstrate that DQT outperforms existing state-of-the-art testing approaches in both code coverage and fault detection, particularly for large, complex apps. The faults detected by DQT have been reproduced and reported to developers; so far, 21 of the reported issues have been explicitly confirmed, and 14 have been fixed. Yuanhong Lan, Minxue Pan, Wenhua Yang 0001, Tian Zhang 0001, Xuandong Li |
ICSE | 5 |
| 2024 | Beyond Manual Modeling: Automating GUI Model Generation Using Design DocumentsabstractGUI models encapsulate the desired visual appearance and interactive behaviors of applications, facilitating various downstream tasks like model-based testing (MBT). Manually constructing high-quality GUI models is not only labor-intensive and costly but also prone to errors, particularly as applications evolve and require frequent model updates. Existing automated approaches for GUI model generation heavily rely on reverse engineering, where the models are abstractions of the code. As a result, they are not suitable for MBT to test functional issues because they are consistent with the code. Meanwhile, valuable development artifacts such as UI/UX design documents, which reflect design intentions, are often overlooked. In this paper, a novel approach named DemGen is proposed to seek a unique pathway for GUI model generation. Leveraging design documents, DemGen employs computer vision pre-trained models in conjunction with a rule-based correction mechanism to identify GUI elements and their intended behaviors as defined in those documents. Subsequently, the identified content is transformed into a formal GUI model adhering to the IFML modeling language. Our evaluation, conducted in collaboration with an industry partner on commercial applications, demonstrates the effectiveness and efficiency of DemGen in GUI element recognition and GUI model generation. Moreover, we conducted a comparative analysis of manual, automated, and hybrid modeling techniques, assessing the usefulness of generated models on MBT tasks. Shaoheng Cao, Renyi Chen, Minxue Pan, Wenhua Yang 0001, Xuandong Li |
ASE | 4 |
| 2024 | An Empirical Study on Python Library Dependency and Conflict IssuesabstractWith the rapid development of open-source communities, code reuse in Python projects is increasingly common. Developers heavily rely on third-party libraries from the Python central repository. They need to write specific configuration scripts with version constraints to ensure the correct version of dependent libraries when building projects. However, existing research focuses on direct dependency libraries and ignores potential dependencies that may exist in other dependency configuration files (e.g., requirements-dev.txt for development dependencies). To fill this gap, we conduct an in-depth comprehensive study to quantify the distribution of direct and potential dependencies, correlation, and classification along with detection tools of dependency conflict issues with 278 top popular Python library projects, which were collected from the prominent dependency tracking system Libraries.io. Specifically, we first investigate the magnitude distribution of dependencies by parsing library source files. Second, we visualize dependencies among Python libraries to determine the correlation. We then classify types of dependency conflicts from the perspective of third-party libraries. Finally, we compare Python dependency conflict detection and resolution tools for researchers and developers. Our findings show that third-party libraries containing dependencies are more common, with 79.1% having at least one dependent library. Moreover, the dependency relationship among Python libraries is intricate, which generates lots of dependency conflict issues such as version conflicts. The main cause of issues is the conflict among third-party libraries, accounting for 60.13%. Our findings can help developers better understand library dependencies and provide them insights on how to better manage them. Yu Zhou 0010, Yasir Hussain, Wenhua Yang 0001 |
QRS | 4 |
| 2024 | How accessibility affects other quality attributes of software? A case study of GitHub
Yaxin Zhao, Lina Gong, Wenhua Yang 0001, Yu Zhou 0010 |
Sci. Comput. Program. | 3 |
| 2024 | Richen: Automated enrichment of Git documentation with usage examples and scenariosabstractAbstract As the predominant modern version control system, Git has become an indispensable tool for both commercial and open‐source software projects. It substantially improves software development effectiveness and efficiency through its distributed version control system, fostering seamless collaboration among teams and across locations. However, research has found that many developers have doubts about using Git commands, while the official Git documentation is rather scanty, that is, lacking sufficient explanations and examples. To help developers learn and use Git commands, we propose the first approach (Richen) for enriching Git documentation with usage examples and scenarios by leveraging crowd knowledge from Stack Overflow. Richen retrieves Git‐related posts from Stack Overflow, extracts relevant Q&A pairs, and selects representative command usages, including usage examples and scenarios, for different Git commands. Experimental results have shown that Richen can extract informative and concise command usages for Git commands. Compared with alternative methods adapted from API usage mining, the command usages obtained by Richen have significant advantages in terms of relevance, readability, and usability. Furthermore, we have shown through an empirical study that the command usages extracted by Richen can better help developers complete Git command‐related tasks. Chaochao Shen, Wenhua Yang 0001, Haitao Jia, Minxue Pan, Yu Zhou 0010 |
J. Softw. Evol. Process. | 2 |
| 2024 | How Important Are Good Method Names in Neural Code Generation? A Model Robustness PerspectiveabstractPre-trained code generation models (PCGMs) have been widely applied in neural code generation, which can generate executable code from functional descriptions in natural languages, possibly together with signatures. Despite substantial performance improvement of PCGMs, the role of method names in neural code generation has not been thoroughly investigated. In this article, we study and demonstrate the potential of benefiting from method names to enhance the performance of PCGMs from a model robustness perspective. Specifically, we propose a novel approach, named neu RA l co D e gener A tor R obustifier (RADAR). RADAR consists of two components: RADAR -Attack and RADAR -Defense. The former attacks a PCGM by generating adversarial method names as part of the input, which are semantic and visual similar to the original input but may trick the PCGM to generate completely unrelated code snippets. As a countermeasure to such attacks, RADAR -Defense synthesizes a new method name from the functional description and supplies it to the PCGM. Evaluation results show that RADAR -Attack can reduce the CodeBLEU of generated code by 19.72% to 38.74% in three state-of-the-art PCGMs (i.e., CodeGPT, PLBART, and CodeT5) in the fine-tuning code generation task and reduce the Pass@1 of generated code by 32.28% to 44.42% in three state-of-the-art PCGMs (i.e., Replit, CodeGen, and CodeT5+) in the zero-shot code generation task. Moreover, RADAR -Defense is able to reinstate the performance of PCGMs with synthesized method names. These results highlight the importance of good method names in neural code generation and implicate the benefits of studying model robustness in software engineering. Guang Yang 0019, Yu Zhou 0010, Wenhua Yang 0001, Tao Yue 0002, Xiang Chen 0005, Taolue Chen 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2023 | Understanding and Enhancing Issue Prioritization in GitHubabstractGitHub has become a prominent platform for open source software development, facilitating collaboration and communication among a diverse group of contributors. Efficient issue tracking is a crucial aspect of managing projects on GitHub, and labels serve as one of the primary mechanisms for issue prioritization, while various other issue features are also utilized by issue handlers for the same purpose. However, in large projects, prioritizing issues remains a challenge, and the efficacy of using labels or other issue features for prioritization is not well understood. To address this knowledge gap, we conduct a comprehensive empirical study that investigates the role of labels in GitHub issue prioritization, examines the influence of various issue features on prioritization, and assesses the performance of different ranking algorithms based on these impactful features. Our study, conducted on a dataset comprising data from over 1.5 million issues across diverse GitHub projects, provides valuable insights for issue handling in open source platforms and offers guidance for future research in this domain. Specifically, the study reveals the limited effectiveness of labels in issue prioritization, highlights the significance of certain issue features in the prioritization process, and compares the performance of various ranking algorithms for issue prioritization to support issue handlers. Yingying He, Wenhua Yang 0001, Minxue Pan, Yasir Hussain, Yu Zhou 0010 |
ASE | 2 |
| 2023 | VALAR: Streamlining Alarm Ranking in Static Analysis with Value-Flow Assisted Active LearningabstractStatic analyzers play a critical role in program defects and security vulnerabilities detection. Despite their importance, the widespread adoption of static analysis techniques in industrial development faces numerous obstacles, among which the high rate of false alarms constitutes a significant one. To address this issue, we propose a novel approach called Valar, which performs alarm ranking for advanced value-flow analysis using the active learning technique. Active learning algorithms minimize the manual effort for alarm inspection by maximizing the effect of each user labeling in recognizing true/false alarms. Meanwhile, the value-flows provide Valar with a concise and comprehensive summary of the operational semantics about programs. Based on this, Valar is able to reason about the potential correlations between alarms and prioritize the most profitable unlabeled alarm. Additionally, the accuracy of Valar increases as more user labels are given and Valar's active learning model is further refined. We evaluate Valar on 20 real-world C/C++ programs using three value-flow based checkers. Our experimental results demonstrated that Valar significantly lowers the priorities of false alarms with most true alarms ranked high. Notably, Valar ranked all true alarms in the top 47% in 90% projects and ranked 90% true alarms in the top 22% in 75% projects. Furthermore, Valar has no requirement for pretraining and has a negligible computation time of less than 0.1s for each alarm prioritization. Wenhua Yang 0001, Minxue Pan |
ASE | 3 |
| 2023 | Mobile Test Script Generation from Natural Language DescriptionsabstractMobile applications are increasingly integral to our daily lives. Currently, the correctness of GUI functions of mobile application is mainly ensured by executing manually written test scripts. However, manually writing these test scripts is not only time-consuming but also costly. Moreover, test scripts are highly vulnerable to application modifications and prone to corruption. In this paper, we propose a novel approach for writing test scripts that enables testers to directly express test intents in natural language within the script. Additionally, we present a new test script generation tool that transforms these test intents into their corresponding test events. Our proposed tool, named GenDroid, employs pre-trained models in conjunction with random forest to facilitate the conversion of test intents into the respective test scripts. To further alleviate the workload of testers and enable them to focus on composing critical test intents, we leverage the application’s UI transfer graph to facilitate the automated generation of other test events, such as jump actions, throughout the generation process. Our results indicate an intent coverage of 88.1%, a notable 20.68% improvement compared to the similar-purpose tool, seq2act. Wenhua Yang 0001, Minxue Pan |
QRS | 4 |
| 2023 | Git Merge Conflict Resolution Leveraging Strategy Classification and LLMabstractIn the realm of collaborative software development, version control systems (VCS) like Git play an indispensable role, enabling concurrent development and facilitating seamless integration of disparate code contributions. Despite these benefits, merge conflicts resulting from simultaneous changes to identical code lines often pose significant challenges to the integration process. Addressing this challenge, our paper introduces a novel two-stage approach, termed as CHATMERGE, for resolving Git merge conflicts. CHATMERGE pioneers a unique strategy that employs machine learning to initially predict resolution strategies, and subsequently leverages a large language model, ChatGPT, to create resolutions for conflicts that necessitate complex resolution strategies. A series of comprehensive experiments validate CHATMERGE’s efficacy, demonstrating its impressive alignment with historical manual resolutions and its superior performance relative to existing, publicly accessible tools. The paper further explores the influence of various classification algorithms and the prompt construction process for ChatGPT, providing further insights into the merge conflict resolution process. Moreover, to foster continued advancements in this area, CHATMERGE, along with its associated training and testing datasets, is made publicly available, offering a valuable resource for both developers and researchers. This work, therefore, provides both an innovative solution to merge conflict resolution and a strong foundation for future explorations in this domain. Chaochao Shen, Wenhua Yang 0001, Minxue Pan, Yu Zhou 0010 |
QRS | 2 |
| 2023 | Understanding the Topics and Challenges of GPU Programming by Classifying and Analyzing Stack Overflow PostsabstractGPUs have cemented their position in computer systems, not restricted to graphics but also extensively used for general-purpose computing. With this comes a rapidly expanding population of developers using GPUs for programming. However, programming with GPUs is notoriously difficult due to their unique architecture and constant evolution. A large number of developers have encountered problems of one kind or another, and many of them have turned to Q&A sites for help. Unfortunately, there has been no prior work to comprehensively study the topics discussed and challenges encountered by developers in GPU programming. To fill this knowledge gap, we conduct a comprehensive study to understand the topics and challenges of GPU programming using Stack Overflow. We collect 25,269 relevant posts from Stack Overflow, propose a novel approach that combines automatic techniques and manual thematic analysis to extract topics, and build a taxonomy of topics with detailed discussions of the popularity, difficulty, and changing trends of these topics. In addition, we analyzed relevant posts through extensive manual efforts to understand the challenges of each topic and to summarize them for future research. Wenhua Yang 0001, Minxue Pan |
ESEC/SIGSOFT FSE | 1 |
| 2023 | Understanding the Role of Stack Overflow in Supporting Software Development Tasks: A Research PerspectiveabstractStack Overflow is a Q&A website that is popular among developers and extensively used in software engineering (SE) research. A significant body of research has examined how Stack Overflow can assist with software development tasks, such as recommending APIs. However, while researchers have recognized the importance of Stack Overflow in SE research related to software development tasks, the specific ways in which it is utilized and the reasons for its widespread usage in research have not been thoroughly explored. To address these knowledge gaps, we conducted the first study to understand the role of Stack Overflow in assisting with SE research regarding software development tasks by systematically examining relevant and high-quality research works. Meanwhile, we carried out a qualitative survey to gain insight into why researchers choose to utilize Stack Overflow in SE research and to solicit suggestions for the better use of Stack Overflow in research. The study identifies trends in the research area, prominent researchers and organizations, and the types of tasks that utilize Stack Overflow in research, with coding and debugging being the most common. Moreover, it examines how Stack Overflow data is utilized in SE research regarding software development tasks, including searching, training models, and mining associations. Our qualitative survey of researchers indicates that the popularity of Stack Overflow stems from its comprehensive explanations of technical topics that are often not found in documentation or manuals. The findings provide a comprehensive understanding of the role of Stack Overflow in SE research regarding software development tasks, and offer actionable implications for both researchers and stakeholders of Stack Overflow to facilitate future research and improvements. Wenhua Yang 0001, Chaochao Shen |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2023 | Git command recommendations using crowd-sourced knowledge
Haitao Jia, Wenhua Yang 0001, Chaochao Shen, Minxue Pan, Yu Zhou 0010 |
Inf. Softw. Technol. | 2 |
| 2023 | Ensure: Towards Reliable Control of Cyber-Physical Systems Under UncertaintyabstractCyber-physical systems (CPSs) are complex ensembles of physical and cyber components that cooperate to offer dynamic and adaptive functionalities. Uncertainty can arise from a plethora of sources in the entangled components, ranging from the unreliable perception, the nondeterministic action effects, to even the changes in the environment. Existing controlling approaches, such as those using Markov decision process, have limited ability in handling uncertainty. To address the challenge, in this article, we novelly propose using partially observable Markov decision processes (POMDPs) to model CPS under uncertainty and show that common types of uncertainties can be modeled by partial observations and nondeterministic actions over probabilistic distributions. With POMDPs, strategies that can optimally control CPS are synthesized. We further propose a strategywise verification method, which resolves the difficult problem of verifying the entire POMDP, to offer reliable controlling strategies. Experiments on two representative cases of CPS show promising results compared with existing approaches. Wenhua Yang 0001, Chang Xu 0001, Minxue Pan, Yu Zhou 0010 |
IEEE Trans. Reliab. | 1 |
| 2022 | Sequence-Aware API Recommendation Based on Collaborative FilteringabstractAPI recommendation is crucial to improve programmers’ productivity. A lot of work has been proposed to improve the accuracy of API recommendations. In the existing work, many metrics, such as Precision, Recall, and MAP are used to evaluate the accuracy of the recommendation. These metrics can well reflect the ability to distinguish useful APIs from the candidate set, but they cannot evaluate the ability to determine the priority of useful APIs with each other. The priority between related APIs directly determines whether the recommended results are practical for developers. From this perspective, inspired by the sequence-aware recommendation, this paper constructs an API recommendation method with sequence awareness and designs new metrics to evaluate the method’s ability to determine the priority of useful APIs. The experimental results show that, compared with the baseline, the proposed method not only achieves better results on the common widely-used metrics but also outperforms the baseline method concerning the newly proposed sequence metrics. Yongchao Wang 0003, Yu Zhou 0010, Taolue Chen 0001, Wenhua Yang 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2022 | Meaningful Update and Repair of Markov Decision Processes for Self-Adaptive Systems
Wenhua Yang 0001, Minxue Pan, Yu Zhou 0010 |
J. Comput. Sci. Technol. | 1 |
| 2022 | Automatic source code summarization with graph attention networksabstractSource code summarization aims to generate concise descriptions for code snippets in a natural language, thereby facilitates program comprehension and software maintenance. In this paper, we propose a novel approach– GSCS –to automatically generate summaries for Java methods, which leverages both semantic and structural information of the code snippets. To this end, GSCS utilizes Graph Attention Networks to process the tokenized abstract syntax tree of the program, which employ a multi-head attention mechanism to learn node features in diverse representation sub-spaces, and aggregate features by assigning different weights to its neighbor nodes. GSCS further harnesses an additional RNN-based sequence model to obtain the semantic features and optimizes the structure by combining its output with a transformed embedding layer. We evaluate our approach on two widely-adopted Java datasets; the experiment results confirm that GSCS outperforms the state-of-the-art baselines. Yu Zhou 0010, Juanjuan Shen, Wenhua Yang 0001, Tingting Han 0001, Taolue Chen 0001 |
J. Syst. Softw. | 4 |
| 2022 | Do Developers Really Know How to Use Git Commands? A Large-scale Study Using Stack OverflowabstractGit, a cross-platform and open source distributed version control tool, provides strong support for non-linear development and is capable of handling everything from small to large projects with speed and efficiency. It has become an indispensable tool for millions of software developers and is the de facto standard of version control in software development nowadays. However, despite its widespread use, developers still frequently face difficulties when using various Git commands to manage projects and collaborate. To better help developers use Git, it is necessary to understand the issues and difficulties that they may encounter when using Git. Unfortunately, this problem has not yet been comprehensively studied. To fill this knowledge gap, in this article, we conduct a large-scale study on Stack Overflow, a popular Q&A forum for developers. We extracted and analyzed 80,370 relevant questions from Stack Overflow, and reported the increasing popularity of the Git command questions. By analyzing the questions, we identified the Git commands that are frequently asked and those that are associated with difficult questions on Stack Overflow to help understand the difficulties developers may encounter when using Git commands. In addition, we conducted a survey to understand how developers learn Git commands in practice, showing that self-learning is the primary learning approach. These findings provide a range of actionable implications for researchers, educators, and developers. Wenhua Yang 0001, Minxue Pan, Chang Xu 0001, Yu Zhou 0010 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2021 | Hybrid Collaborative Filtering-Based API RecommendationabstractAutomatic API recommendations can liberate software developers from labor-intensive programming tasks. Collaborative filtering (CF) techniques, which have been proved to be superior to other classic techniques, are widely used in recommendation tasks such as music, book, and goods recommendations, but are rarely used in the recommendation of APIs. In this paper, we employ the hybrid of CF techniques to build an API recommendation system. More precisely, We treat the API recommendation task as an item recommendation problem, where method declarations are regarded as users, API calls are regarded as items. First, we use the memory-based CF technique to find the most similar projects, collect the most similar declarations, and take API calls used by the considered declarations together to generate a rating matrix. Next, we use the model-based CF technique to complete the missing values in the rating matrix, then a ranked list of APIs is generated based on the completed rating matrix and sent to the developers as a recommendation result. Experimental results show that compared with the state-of-the-art work, the proposed approach can achieve better performance in terms of a comprehensive set of metrics, such as Success Rate, Precision, Recall, MRR, and NDCG for the top-1, top-3 and top-5 recommended APIs. Yongchao Wang 0003, Yu Zhou 0010, Taolue Chen 0001, Wenhua Yang 0001 |
QRS | 5 |
| 2021 | Personalized API RecommendationsabstractApplication Programming Interfaces (APIs) play an important role in modern software development. Developers interact with APIs on a daily basis and thus need to learn and memorize those APIs suitable for implementing the required functions. This can be a burden even for experienced developers since there exists a mass of available APIs. API recommendation techniques focus on assisting developers in selecting suitable APIs. However, existing API recommendation techniques have not taken the developers personal characteristics into account. As a result, they cannot provide developers with personalized API recommendation services. Meanwhile, they lack the support for self-defined APIs in the recommendation. To this end, we aim to propose a personalized API recommendation method that considers developers’ differences. Our API recommendation method is based on statistical language. We propose a model structure that combines the N-gram model and the long short-term memory (LSTM) neural network and train predictive models using API invoking sequences extracted from GitHub code repositories. A general language model trained on all sorts of code data is first acquired, based on which two personalized language models that recommend personalized library APIs and self-defined APIs are trained using the code data of the developer who needs personalized services. We evaluate our personalized API recommendation method on real-world developers, and the experimental results show that our approach achieves better accuracy in recommending both library APIs and self-defined APIs compared with the state-of-the-art. The experimental results also confirm the effectiveness of our hybrid model structure and the choice of the LSTM’s size. Wenhua Yang 0001, Yu Zhou 0010 |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2020 | SAT and LP Collaborative Bounded Timing Analysis of Scenario-Based SpecificationsabstractTiming analysis of scenario-based specifications (SBS) such as message sequence charts and UML interaction models plays an essential role in the design phase of real-time system development. However, it is time-consuming and labor-intensive to conduct analysis on the satisfiability of the timing constraints. In this article, we propose a novel SAT and linear programming (LP) collaborative timing analysis approach named TASSAT for SBS. Instead of using depth-first traversal algorithms, TASSAT encodes the structures of the SBS into propositional formulas and use the SAT solver to find candidate paths. The timing analysis of candidate paths is then reduced to LP problems, where irreducible infeasible set of the infeasible path can be used to prune unnecessary search space of the SAT solver. The experimental results show that TASSAT is effective and offers better performance than existing tools in terms of both time consumption and memory footprint. Longlong Lu, Wenhua Yang 0001, Minxue Pan, Tian Zhang 0001 |
Internetware | 2 |
| 2020 | Developer portraying: A quick approach to understanding developers on OSS platforms
Wenhua Yang 0001, Minxue Pan, Yu Zhou 0010 |
Inf. Softw. Technol. | 1 |
| 2019 | Extracting Mapping Relations for Mobile User Interface TransformationabstractThe development of mobile apps has become the current mantra for any business' success. The rise of many types of mobile devices and mobile OS has instantly created the need to develop multiple versions for the same app. In order to grasp as much market share as possible, it is desirable to have all the versions of an app demonstrate similar user interface (UI) appearances, to make users feel comfortable when switching from one platform to another and more likely to stick to the app. However, to ensure consistent UIs among cross-platform versions can be a challenging and costly endeavor, since different platforms have their own UI controls and programming languages. In this paper, we propose an automatic approach to transforming mobile app UIs across platforms, and illustrate our approach by transforming the UIs of iOS apps to Android ones. We leverage the enormous existing apps carefully designed by developers to achieve similar UI effects between iOS and Android versions, since these apps contain valuable knowledge of mapping relations between the iOS and Android UI controls. Starting from the reverse engineering of these apps, our approach separates each user interface into modules of adequate sizes. Then it maps the modules from both versions that contribute to the same visual and functional effect, and automatically mines the mapping relations. By applying the mined relations, our approach has successfully transformed the iOS app UIs into Android app UIs, as confirmed by a series of experiments. Ruihua Ji, Junyu Pei, Wenhua Yang 0001, Juan Zhai, Minxue Pan, Tian Zhang 0001 |
Internetware | 3 |
| 2019 | Mining and Comparing User Reviews across Similar Mobile AppsabstractWith the rapid development of the market for mobile apps, there are a number of apps with similar functions. To gain an advantage in such a competitive environment, developers need to understand not only the strengths and weaknesses of their app but also competitive apps. User reviews contain valuable information for comparing similar apps from user preference. In this paper, we propose UISAT (User-review mining via topic Identification, Sentiment Analysis and Topic matching across apps), an automated approach to compare user reviews from similar apps with the goal of mining user feedback from competitive apps by (i) extracting the hidden topics from large volumes of user reviews using topic modeling, (ii) combining a rule-based model, user rating and user-helpful for sentiment analysis of topics and (iii) matching relevant topics across apps. Empirical studies demonstrate that UISAT is effective and promising for developers to build and maintain a more competitive app. Yanqi Su, Yongchao Wang 0003, Wenhua Yang 0001 |
MSN | 3 |
| 2019 | Augmenting Java method comments generation with context information based on neural networks
Yu Zhou 0010, Wenhua Yang 0001, Taolue Chen 0001 |
J. Syst. Softw. | 3 |
| 2018 | Efficient validation of self-adaptive applications by counterexample probability maximization
Wenhua Yang 0001, Chang Xu 0001, Minxue Pan, Chun Cao, Xiaoxing Ma, Jian Lu 0001 |
J. Syst. Softw. | 1 |
| 2018 | Improving Verification Accuracy of CPS by Modeling and Calibrating Interaction UncertaintyabstractCyber-Physical Systems (CPS) intrinsically combine hardware and physical systems with software and network, which are together creating complex and correlated interactions. CPS applications often experience uncertainty in interacting with environment through unreliable sensors. They can be faulty and exhibit runtime errors if developers have not considered environmental interaction uncertainty adequately. Existing work in verifying CPS applications ignores interaction uncertainty and thus may overlook uncertainty-related faults. To improve verification accuracy, in this article we propose a novel approach to verifying CPS applications with explicit modeling of uncertainty arisen in the interaction between them and the environment. Our approach builds an Interactive State Machine network for a CPS application and models interaction uncertainty by error ranges and distributions. Then it encodes both the application and uncertainty models to Satisfiability Modulo Theories (SMT) formula to leverage SMT solvers searching for counterexamples that represent application failures. The precision of uncertainty model can affect the verification results. However, it may be difficult to model interaction uncertainty precisely enough at the beginning, because of the uncontrollable noise of sensors and insufficient data sample size. To further improve the accuracy of the verification results, we propose an approach to identifying and calibrating imprecise uncertainty models. We exploit the inconsistency between the counterexamples’ estimate and actual occurrence probabilities to identify possible imprecision in uncertainty models, and the calibration of imprecise models is to minimize the inconsistency, which is reduced to a Search-Based Software Engineering problem. We experimentally evaluated our verification and calibration approaches with real-world CPS applications, and the experimental results confirmed their effectiveness and efficiency. Wenhua Yang 0001, Chang Xu 0001, Minxue Pan, Xiaoxing Ma, Jian Lu 0001 |
ACM Trans. Internet Techn. | 1 |
| 2016 | Suppressing detection of inconsistency hazards with pattern learning
Chang Xu 0001, Wenhua Yang 0001, Xiaoxing Ma, Ping Yu 0004, Jian Lu 0001 |
Inf. Softw. Technol. | 3 |
| 2015 | A survey on dependability improvement techniques for pervasive computing systems
Wenhua Yang 0001, Yepang Liu 0001, Chang Xu 0001, Shing-Chi Cheung |
Sci. China Inf. Sci. | 1 |
| 2014 | SHAP: Suppressing the Detection of Inconsistency Hazards by Pattern LearningabstractContext-aware applications rely on contexts derived from sensory data to adapt their behavior. However, contexts can be inconsistent and cause application anomaly or crash. One popular solution is to detect and resolve context inconsistencies at runtime. However, we observe that many detected inconsistencies do not indicate real context problems. Instead, they are caused by improper inconsistency detection. These inconsistencies are harmless, and their resolution is unnecessary or may even cause new problems. We name them inconsistency hazards. Inconsistency hazards should be suppressed, but their occurrences resemble normal inconsistencies. In this paper, we present a pattern-learning based approach SHAP to suppressing the detection of inconsistency hazards. Our key insight is that the detection of such hazards is subject to certain patterns of context changes. These patterns, although difficult to specify manually, can be learned effectively from historical inconsistency detection data. We evaluated our SHAP experimentally through three context-aware applications. The results reported that SHAP can automatically suppress the detection of over 90% inconsistency hazards, while preserving the detection of over 98% normal inconsistencies, with only negligible overhead. Chang Xu 0001, Wenhua Yang 0001, Ping Yu 0004, Xiaoxing Ma, Jiang Lu |
APSEC (1) | 3 |
| 2014 | Automated recommendation of dynamic software update points: an exploratory studyabstractDue to the demand for bugs fixing and feature enhancements, developers inevitably need to update in-use software systems. Instead of shutting down a running system before updating, it is often desirable and sometimes mandatory to patch the running software system on the fly, with a mechanism generally referred as dynamic software updating (DSU). Practical DSU strategies often require manual specification of update points in the program for performing dynamic updates. At these points DSU systems will update the program code, and also migrate the program state to the new version program (using transformation functions). However, finding appropriate update points is non-trivial because the choice of update points has great influence on two competing factors: the timeliness of DSU and the complexity of transformation functions; and to strike a good balance between them requires a deep understanding of both versions of the program. In this exploratory paper, we conceive an automated approach to the recommendation of update points for developers. We conduct a set of preliminary experiments with a real world software update case to examine the feasibility of the approach. Xiaoxing Ma, Chang Xu 0001, Wenhua Yang 0001 |
Internetware | 4 |
| 2014 | Verifying self-adaptive applications suffering uncertaintyabstractSelf-adaptive applications address environmental dynamics systematically. They can be faulty and exhibit runtime errors when environmental dynamics are not considered adequately. It becomes more severe when uncertainty exists in their sensing and adaptation to environments. Existing work verifies self-adaptive applications, but does not explicitly consider environmental constraints or uncertainty. This gives rise to inaccurate verification results. In this paper, we address this problem by proposing a novel approach to verifying self-adaptive applications suffering uncertainty in their environmental interactions. It builds Interactive State Machine (ISM) models for such applications and verifies them with explicit consideration of environmental constraints and uncertainty. It then refines verification results by prioritizing counterexamples according to their probabilities. We experimentally evaluated our approach with real-life self-adaptive applications, and the experimental results confirmed its effectiveness. Our approach reported 200-660% more counterexamples than not considering uncertainty, and eliminated all false counterexamples caused by ignoring environmental constraints. Wenhua Yang 0001, Chang Xu 0001, Yepang Liu 0001, Chun Cao, Xiaoxing Ma, Jian Lu 0001 |
ASE | 1 |
| 2013 | Environment rematching: Toward dependability improvement for self-adaptive applicationsabstractSelf-adaptive applications can easily contain faults. Existing approaches detect faults, but can still leave some undetected and manifesting into failures at runtime. In this paper, we study the correlation between occurrences of application failure and those of consistency failure. We propose fixing consistency failure to reduce application failure at runtime. We name this environment rematching, which can systematically reconnect a self-adaptive application to its environment in a consistent way. We also propose enforcing atomicity for application semantics during the rematching to avoid its side effect. We evaluated our approach using 12 self-adaptive robot-car applications by both simulated and real experiments. The experimental results confirmed our approach's effectiveness in improving dependability for all applications by 12.5-52.5%. Chang Xu 0001, Wenhua Yang 0001, Xiaoxing Ma, Chun Cao, Jian Lu 0001 |
ASE | 2 |