Shengcheng Yu

dblp:243/7123 · DBLP profile ↗
← Back
26ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0003-4640-8637ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 24 · 10 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 A survey on large language models for software engineering
abstract
Abstract Software engineering (SE) is the systematic design, development, maintenance, and management of software applications underpinning the digital infrastructure of our modern world. Very recently, the SE community has seen a rapidly increasing number of techniques employing large language models (LLMs) to automate a broad range of SE tasks. Nevertheless, existing information on the applications, effects, and possible limitations of LLMs within SE is still not well-studied. In this paper, we provide a systematic survey to summarize the current state-of-the-art research in the LLM-based SE community. We summarize 62 representative LLMs of Code across three model architectures, 15 pre-training objectives across four categories, and 16 downstream tasks across five categories. We then present a detailed summarization of the recent SE studies for which LLMs are commonly utilized, including 926 studies for 112 specific code-related tasks across five crucial phases within the SE workflow. We also discuss several critical aspects during the integration of LLMs into SE, such as empirical evaluation, benchmarking, security and reliability, domain tuning, compressing, and distillation. Finally, we highlight several challenges and potential opportunities in applying LLMs for future SE studies, such as exploring domain LLMs and constructing clean evaluation datasets. Overall, our work can help researchers gain a comprehensive understanding about the achievements of the existing LLM-based SE studies and promote the practical application of these techniques. Our artifacts are publicly available and will be continuously updated at the living repository https://github.com/iSEngLab/AwesomeLLM4SE .
Quanjun Zhang, Chunrong Fang, Shengcheng Yu, Weisong Sun, Yun Yang 0001, Zhenyu Chen 0001
Sci. China Inf. Sci.5
2026 GUI test migration via LLM with scenario-granularity understanding
Shengcheng Yu, Chunrong Fang, Junyang Xing, Jia Liu 0015, Zhenyu Chen 0001
Frontiers Comput. Sci.2
2026 LLM-based Crowdsourced Test Report Clustering
abstract
The openness of crowdsourced testing introduces diversity in testing results. However, it also leads to a large volume of test reports, many of which highlight the same recurring issues. While these reports provide valuable feedback, their redundancy makes it inefficient for developers to review the reports and identify bugs. Crowdsourced test report clustering has been proposed to mitigate this problem, allowing developers to focus only on the representative reports from each cluster. However, existing methods primarily rely on embedding features extracted from reports for clustering, which limits their ability to generate accurate and interpretable clusters due to a lack of deeper semantic understanding of the reports. To address the aforementioned challenge, we propose LLMCluster , a novel method for crowdsourced test report clustering based on Large Language Models (LLMs). LLMCluster employs an iterative clustering strategy. In each iteration, LLMCluster processes a subset of reports by instructing the LLM to disregard surface-level variations in expression, analyze the core issue in each report, and group reports addressing the same issue into new or existing clusters. After the iterative clustering process, LLMCluster applies correction algorithms to ensure the completeness and validity of the clustering result. Finally, LLMCluster utilizes the LLM to generate concise summaries for each cluster, making the results more intuitive and interpretable. Experimental results show that LLMCluster outperforms state-of-the-art methods across six commonly used clustering evaluation metrics. Additionally, the cluster summaries generated by LLMCluster semantically align well with manually written summaries.
Yuchen Ling, Shengcheng Yu, Chunrong Fang, Zhenyu Chen 0001
ACM Trans. Softw. Eng. Methodol.2
2026 Test Script Intention Generation for Mobile Application via GUI Image and Code Understanding
abstract
Testing is the most direct and effective technique to ensure software quality. Test scripts always play a more important role in mobile app testing than test cases for source code, due to the GUI-intensive and event-driven characteristics of mobile applications (app). Test scripts focus on user interactions and the corresponding response events, which is significant for testing the target app functionalities. Therefore, it is critical to understand the test scripts for better script maintenance and modification. There exist some mature code understanding (i.e., code comment generation, code summarization) technologies that can be directly applied to functionality source code with business logic. However, such technologies will have difficulties when being applied to test scripts, because test scripts are loosely linked to Apps under Test (AUT) by widget selectors, and do not contain business logic themselves. In order to solve the test script understanding gap, this article presents a novel approach, namely TestIntention , to infer the intention of GUI test scripts. Test intention refers to the user expectations of app behaviors for specific operations . TestIntention formalizes test scripts with an operation sequence model. For each operation within the sequence, TestIntention extracts the target widget selector and links the selector to the GUI layout information or the corresponding response events. For widgets identified by XPath , TestIntention utilizes the image understanding technologies to explore the detailed information of the widget images, the intention of which is understood with a deep learning model. For widgets identified by ID , TestIntention first maps the selectors to the response methods with business logic, and then adopts code understanding technologies to describe code in natural language form. Results of all operations are combined to generate test intention for test scripts. An empirical experiment including different metrics proves the outstanding performance of TestIntention , outperforming baselines by much. Also, it is shown that TestIntention can save about 80% developers’ time to understand test scripts.
Shengcheng Yu, Chunrong Fang, Jia Liu 0015, Zhenyu Chen 0001
ACM Trans. Softw. Eng. Methodol.1
2026 Human Cognitive Pattern Simulation for Crowdsourced Test Report Consistency Detection
abstract
Crowdsourced testing has emerged as a prominent paradigm in software testing by leveraging the diversity of crowdworkers. In this paradigm, crowd-workers are required to submit a test report for each identified bug, which typically contains a textual description and a bug screenshot. However, due to varying worker expertise, many reports exhibit inconsistencies between the textual description and the bug screenshot, which hinder the report review process. Existing methods address this issue by automatically detecting report consistency, typically through matching the UI widgets referenced in the textual description with those visible in the bug screenshot. However, such methods focus only on surface-level element correspondence and fail to capture the abstract bug semantics, such as the functional meaning and bug-triggering context. Consequently, they lack the ability to detect more subtle but realistic inconsistencies. To bridge this gap, we propose INCONHUNTER, a novel method for crowdsourced test report consistency detection that explicitly simulates human cognitive pattern. In this pattern, humans typically adopt two complementary reasoning strategies. If the textual description allows them to form an expectation about the visual bug features, they assess consistency by verifying if the expected features appear in the bug screenshot. Otherwise, they shift to reasoning about if the bug-triggering context described in the report aligns with the app state shown in the bug screenshot. INCONHUNTERinstantiates this cognitive pattern through two LLM-powered modules, each dedicated to one reasoning strategy. We evaluate INCONHUNTERthrough experiments on our dataset with 2,310 labeled crowdsourced test reports, and results show that INCONHUNTERoutperforms baselines by 14.00%–19.28%, demonstrating superior effectiveness, monetary-based cost efficiency, and alignment with human cognitive pattern.
Yuchen Ling, Shengcheng Yu, Shuguang Chen, Liuming Wang, Chunrong Fang, Jia Liu 0015, Zhenyu Chen 0001
IEEE Trans. Software Eng.2
2025 Multi-dimensional Assessment of Crowdsourced Testing Reports via LLMs
abstract
Crowdsourced testing can markedly enhance test coverage and the discovery rate of potential defects compared to traditional software testing, making it increasingly popular. However, with the widespread use of crowdsourced testing, more and more crowdworkers from various backgrounds are submitting a large number of testing reports to crowdsourced testing platforms, which hinders developers from effectively reviewing the reports. Facing a vast amount of reports with varying quality, manual review is not only time-consuming and labor-intensive but also increases costs. Therefore, how to efficiently review crowdsourced testing reports has become a major challenge. To address this challenge, we propose a multi-dimensional assessment method for crowdsourced testing reports based on large language models. This method not only inherits the textuality dimension widely used in traditional report assessment but also innovatively introduces two new dimensions: adequacy and competitiveness. It comprehensively assesses the quality of crowdsourced testing reports from multiple perspectives, aiming to better screen for high-quality crowdsourced testing reports. Through experimental analysis conducted on three different applications, we have proven the consistency of our method with human raters across various dimensions, and we have also observed an enhancement in the efficiency of report assessment.
Shengcheng Yu, Zhenyu Chen 0001
ASE3
2025 Redefining crowdsourced test report prioritization: An innovative approach with large language model
Yuchen Ling, Shengcheng Yu, Chunrong Fang, Guobin Pan, Jia Liu 0008
Inf. Softw. Technol.2
2025 Enhanced Crowdsourced Test Report Prioritization via Image-and-Text Semantic Understanding and Feature Integration
abstract
Crowdsourced testing has gained prominence in the field of software testing due to its ability to effectively address the challenges posed by the fragmentation problem in mobile app testing. The inherent openness of crowdsourced testing brings diversity to the testing outcome. However, it also presents challenges for app developers in inspecting a substantial quantity of test reports. To help app developers inspect the bugs in crowdsourced test reports as early as possible, crowdsourced test report prioritization has emerged as an effective technology by establishing a systematic optimal report inspecting sequence. Nevertheless, crowdsourced test reports consist of app screenshots and textual descriptions, but current prioritization approaches mostly rely on textual descriptions, and some may add vectorized image features at the image-as-a-whole level or widget level. They still lack precision in accurately characterizing the distinctive features of crowdsourced test reports. In terms of prioritization strategy, prevailing approaches adopt simple prioritization based on features combined merely using weighted coefficients, without adequately considering the semantics, which may result in biased and ineffective outcomes. In this paper, we proposeEncrePrior, an enhanced crowdsourced test report prioritization approach via image-and-text semantic understanding and feature integration.EncrePriorextracts distinctive features from crowdsourced test reports. For app screenshots,EncrePriorconsiders the structure (i.e., GUI layout) and the contents (i.e., GUI widgets), viewing the app screenshot from the macroscopic and microscopic perspectives, respectively. For textual descriptions,EncrePriorconsiders the Bug Description and Reproduction Step as the bug context. During the prioritization, we do not directly merge the features with weights to guide the prioritization. Instead, in order to comprehensively consider the semantics, we adopt a prioritize-reprioritize strategy. This practice combines different features together by considering their individual ranks. The reports are first prioritized on four features separately. Then, the ranks on four sequences are used to lexicographically reprioritize the test reports with an integration of features from app screenshots and textual descriptions. Results of an empirical study show thatEncrePrioroutperforms the representative baseline approachDeepPriorby 15.61% on average, ranging from 2.99% to 63.64% on different apps, and the novelly proposed features and prioritization strategy all contribute to the excellent performance ofEncrePrior.
Chunrong Fang, Shengcheng Yu, Quanjun Zhang, Xin Li 0034, Yulei Liu, Zhenyu Chen 0001
IEEE Trans. Software Eng.2
2025 Improving Retrieval-Augmented Deep Assertion Generation via Joint Training
abstract
Unit testing attempts to validate the correctness of basic units of the software system under test and has a crucial role in software development and testing. However, testing experts have to spend a huge amount of effort to write unit test cases manually. Very recent work proposes a retrieve-and-edit approach to automatically generate unit test oracles,i.e.,assertions. Despite being promising, it is still far from perfect due to some limitations, such as splitting assertion retrieval and generation into two separate components without benefiting each other. In this paper, we propose AG-RAG, a retrieval-augmented automated assertion generation (AG) approach that leverages external codebases and joint training to address various technical limitations of prior work. Inspired by the plastic surgery hypothesis, AG-RAG attempts to combine relevant unit tests and advanced pre-trained language models (PLMs) with retrieval-augmented fine-tuning. The key insight of AG-RAG is to simultaneously optimize the retriever and the generator as a whole pipeline with a joint training strategy, enabling them to learn from each other. Particularly, AG-RAG builds a dense retriever to search for relevant test-assert pairs (TAPs) with semantic matching and a retrieval-augmented generator to synthesize accurate assertions with the focal-test and retrieved TAPs as input. Besides, AG-RAG leverages a code-aware language model CodeT5 as the cornerstone to facilitate both assertion retrieval and generation tasks. Furthermore, AG-RAG designs a joint training strategy that allows the retriever to learn from the feedback provided by the generator. This unified design fully adapts both components specifically for retrieving more useful TAPs, thereby generating accurate assertions. AG-RAG is a generic framework that can be adapted to various off-the-shelf PLMs. We extensively evaluate AG-RAG against six state-of-the-art AG approaches on two benchmarks and three metrics. Experimental results show that AG-RAG significantly outperforms previous AG approaches on all benchmarks and metrics,e.g.,improving the most recent baselineEditASby 20.82% and 26.98% in terms of accuracy. AG-RAG also correctly generates 1739 and 2866 unique assertions that all baselines fail to generate, 3.45X and 9.20X more thanEditAS. We further demonstrate the positive contribution of our joint training strategy,e.g.,AG-RAG improving a variant without the retriever by an average accuracy of 14.11%. Besides, adopting other PLMs can provide substantial advancement,e.g.,AG-RAG with four different PLMs improving EditAS by an average accuracy of 9.02%, highlighting the generalizability of our framework. Overall, our work demonstrates the promising potential of jointly fine-tuning the PLM-based retriever and generator to predict accurate assertions by incorporating external knowledge sources, thereby reducing the manual efforts of unit testing experts in practical scenarios.
Quanjun Zhang, Chunrong Fang, Ruixiang Qian, Shengcheng Yu, Yuan Zhao 0010, Yun Yang 0001, Tao Zheng 0005, Zhenyu Chen 0001
IEEE Trans. Software Eng.5
2024 Practical Non-Intrusive GUI Exploration Testing with Visual-based Robotic Arms
abstract
Graphical User Interface (GUI) testing has been a significant topic in the software engineering community. Most existing GUI testing frameworks are intrusive and can only support some specific platforms, which are quite limited. With the development of distinct scenarios, diverse embedded systems or customized operating systems on different devices do not support existing intrusive GUI testing frameworks. Some approaches adopt robotic arms to replace the interface invoking of mobile apps under test and use computer vision technologies to identify GUI elements. However, some challenges remain unsolved with such approaches. First, existing approaches assume that GUI screens are fixed so that they cannot be adapted to diverse systems with different screen conditions. Second, existing approaches use XY-plane robotic arm system, which cannot flexibly simulate human testing operations. Third, existing approaches ignore the compatibility bugs of apps and only focus on the crash bugs. To sum up, a more practical approach is required for the non-intrusive scenario.
Shengcheng Yu, Chunrong Fang, Mingzhe Du, Yuchen Ling, Zhenyu Chen 0001, Zhendong Su 0001
ICSE1
2024 Effective, Platform-Independent GUI Testing via Image Embedding and Reinforcement Learning
abstract
Software applications (apps) have been playing an increasingly important role in various aspects of society. In particular, mobile apps and web apps are the most prevalent among all applications and are widely used in various industries as well as in people’s daily lives. To help ensure mobile and web app quality, many approaches have been introduced to improve app GUI testing via automated exploration, including random testing, model-based testing, learning-based testing, and so on. Despite the extensive effort, existing approaches are still limited in reaching high code coverage, constructing high-quality models, and being generally applicable. Reinforcement learning-based approaches, as a group of representative and advanced approaches for automated GUI exploration testing, are faced with difficult challenges, including effective app state abstraction, reward function design, and so on. Moreover, they heavily depend on the specific execution platforms (i.e., Android or Web), thus leading to poor generalizability and being unable to adapt to different platforms. This work specifically tackles these challenges based on the high-level observation that apps from distinct platforms share commonalities in GUI design. Indeed, we propose PIRLTest , an effective platform-independent approach for app testing. Specifically, PIRLTest utilizes computer vision and reinforcement learning techniques in a novel, synergistic manner for automated testing. It extracts the GUI widgets from GUI pages and characterizes the corresponding GUI layouts, embedding the GUI pages as states. The app GUI state combines the macroscopic perspective (app GUI layout) and the microscopic perspective (app GUI widget) and attaches the critical semantic information from GUI images. This enables PIRLTest to be platform-independent and makes the testing approach generally applicable on different platforms. PIRLTest explores apps with the guidance of a curiosity-driven strategy, which uses a Q-network to estimate the values of specific state-action pairs to encourage more exploration in uncovered pages without platform dependency. The exploration will be assigned with rewards for all actions, which are designed considering both the app GUI states and the concrete widgets, to help the framework explore more uncovered pages. We conduct an empirical study on 20 mobile apps and 5 web apps, and the results show that PIRLTest is zero-cost when being adapted to different platforms, and can perform better than the baselines, covering 6.3–41.4% more code on mobile apps and 1.5–51.1% more code on web apps. PIRLTest is capable of detecting 128 unique bugs on mobile and web apps, including 100 bugs that cannot be detected by the baselines.
Shengcheng Yu, Chunrong Fang, Xin Li 0034, Yuchen Ling, Zhenyu Chen 0001, Zhendong Su 0001
ACM Trans. Softw. Eng. Methodol.1
2024 Practical, Automated Scenario-Based Mobile App Testing
abstract
The importance of mobile application (app) quality assurance is increasing with the rapid development of the mobile Internet. Automated test generation approaches, as a dominant direction of app quality assurance, follow specific models or strategies, targeting at optimizing the code coverage. Such approaches lead to a huge gap between testing execution and app business logic. Test scripts developed by human testers consider business logic by focusing on testing scenarios. Due to the GUI-intensive feature of mobile apps, human testers always understand app GUI to organize test scripts for scenarios. This inspires us to utilize domain knowledge from app GUI understanding for scenario-based test generation. In this paper, we propose a novel approach,ScenTest, for scenario-based mobile app testing with event knowledge graph (EKG) via GUI image understanding.ScenTesttries to start automated testing by imitating human practices and integrating domain knowledge into scenario-based mobile app testing, realizing fully automated testing on target testing scenarios for the first time.ScenTestextracts four kinds of entities and five kinds of corresponding relationships from crowdsourced test reports, where the test events and app GUI information are presented, and constructs the EKGs for specific scenarios. Then,ScenTestconducts test generation for specific scenarios on different apps with the guidance of EKG with the combination consideration of app current state and testing context. We conduct an evaluation onScenTeston different aspects. The results show that the test generation ofScenTeston the basis of EKG is effective, andScenTestreveals 150+ distinct real-world bugs in specific scenarios compared with representative baselines.
Shengcheng Yu, Chunrong Fang, Mingzhe Du, Zimin Ding, Zhenyu Chen 0001, Zhendong Su 0001
IEEE Trans. Software Eng.1
2023 LLM for Test Script Generation and Migration: Challenges, Capabilities, and Opportunities
abstract
This paper investigates the application of large language models (LLM) in the domain of mobile application test script generation. Test script generation is a vital component of software testing, enabling efficient and reliable automation of repetitive test tasks. However, existing generation approaches often encounter limitations, such as difficulties in accurately capturing and reproducing test scripts across diverse devices, platforms, and applications. These challenges arise due to differences in screen sizes, input modalities, platform behaviors, API inconsistencies, and application architectures. Overcoming these limitations is crucial for achieving robust and comprehensive test automation.By leveraging the capabilities of LLMs, we aim to address these challenges and explore its potential as a versatile tool for test automation. We investigate how well LLMs can adapt to diverse devices and systems while accurately capturing and generating test scripts. Additionally, we evaluate its cross-platform generation capabilities by assessing its ability to handle operating system variations and platform-specific behaviors. Furthermore, we explore the application of LLMs in cross-app migration, where it generates test scripts across different applications and software environments based on existing scripts.Throughout the investigation, we analyze its adaptability to various user interfaces, app architectures, and interaction patterns, ensuring accurate script generation and compatibility. The findings of this research contribute to the understanding of LLMs’ capabilities in test automation. Ultimately, this research aims to enhance software testing practices, empowering app developers to achieve higher levels of software quality and development efficiency.
Shengcheng Yu, Chunrong Fang, Yuchen Ling, Chentian Wu, Zhenyu Chen 0001
QRS1
2023 Test Report Generation for Android App Testing Via Heterogeneous Data Analysis
abstract
The rising of the Android market demands higher quality assurance of Android applications (apps) to sharpen the competitive edge, and techniques for traditional software have problems adapting for mobile apps. Android apps often require testing on a large-scale device cluster, which produces a large amount of test reports consisting of heterogeneous data, e.g., hardware information, GUI screenshots, runtime logs. Such data are hard to merge to be unified analyzed, while they serve as an essential basis for bug inspection and fixing. Existing test report generation or analysis techniques can only handle testing data from different devices separately. They simply list all the information to app developers and have no further processing to summarize test reports. Besides, they neglect the inner connection of the heterogeneous data. Such techniques cannot improve the report reviewing effectiveness and efficiency, and they can hardly find the inner links and rules of the bug occurrence on different devices. As a result, developers still need to devote many efforts to inspect and fix bugs. In this paper, a large amount of test reports are investigated by the authors, as to construct a structured bug model to analyze heterogeneous data of the testing results. According to the investigation, we also define theBug Inconsistencyof testing results from multiple devices and build a novel bug taxonomy. In general, an automated approach is proposed to generate structured and comprehensible test reports from raw testing results from multiple devices. Based on the approach, a tool, namelyBreGat, is implemented to evaluate the classification and deduplication capability of our approach. The experimental results of 30 Android apps on 20 devices show thatBreGatcan successfully cover 83% bug categories and exclude 76% duplicate bugs. Furthermore, a user study involving 16 developers shows that our test reports are more comprehensible andBreGatgreatly improves the bug inspection efficiency compared to the state-of-the-art tool.
Chunrong Fang, Shengcheng Yu, Ting Su 0001, Yuanhan Tian, Yang Liu 0003
IEEE Trans. Software Eng.2
2023 Leveraging Android Automated Testing to Assist Crowdsourced Testing
abstract
Crowdsourced testing is an emerging trend in mobile application testing. The openness of crowdsourced testing provides a promising way to conduct large-scale and user-oriented testing scenarios on various mobile devices, while it also brings a problem, i.e., crowdworkers with different levels of testing experience severely threaten the quality of crowdsourced testing. Currently, many approaches have been proposed and studied to improve crowdsourced testing. However, these approaches do not fundamentally improve the ability of crowdworkers. In essence, the low-quality crowdsourced testing is caused by crowdworkers who are unfamiliar with the App Under Test (AUT) and do not know which part of the AUT should be tested. To address this problem, we propose a testing assistance approach, which leverages Android automated testing (i.e., dynamic and static analysis) to improve crowdsourced testing. Our approach constructs an Annotated Window Transition Graph (AWTG) model for the AUT by merging dynamic and static analysis results. Based on the AWTG model, our approach implements a testing assistance pipeline that provides the test task extraction, test task recommendation, and test task guidance to assist crowdworkers in testing the AUT. We experimentally evaluate our approach on real-world AUTs. The quantitative results demonstrate that our approach can effectively and efficiently assist crowdsourced testing. Besides, the qualitative results from a user study confirm the usefulness of our approach.
Xiuting Ge, Shengcheng Yu, Chunrong Fang
IEEE Trans. Software Eng.2
2023 Mobile App Crowdsourced Test Report Consistency Detection via Deep Image-and-Text Fusion Understanding
abstract
Crowdsourced testing, as a distinct testing paradigm, has attracted much attention in software testing, especially in mobile application (app) testing field. Compared with in-house testing, crowdsourced testing shows superiority with the diverse testing environments when faced with the mobile testing fragmentation problem. However, crowdsourced testing also encounters the low-quality test report problem caused by unprofessional crowdworkers involved with different expertise. In order to handle the submitted reports of uneven quality, app developers have to distinguish high-quality reports from low-quality ones to help the bug inspection. One kind of typical low-quality test report is inconsistent test reports, which means the textual descriptions are not focusing on the attached bug-occurring screenshots. According to our empirical survey, only 18.07% crowdsourced test reports are consistent. Inconsistent reports cause waste on mobile app testing. To solve the inconsistency problem, we propose RECODE to detect the consistency of crowdsourced test reports via deep image-and-text fusion understanding. RECODE is a two-stage approach that first classifies the reports based on textual descriptions into different categories according to the bug feature. In the second stage, RECODE has a deep understanding of the GUI image features of the app screenshots and then applies different strategies to handle different types of bugs to detect the consistency of the crowdsourced test reports. We conduct an experiment on a dataset with over 22k test reports to evaluate RECODE, and the results show the effectiveness of RECODE in detecting the consistency of crowdsourced test reports. Besides, a user study is conducted to prove the practical value of RECODE in effectively helping app developers improve the efficiency of reviewing the crowdsourced test reports.
Shengcheng Yu, Chunrong Fang, Quanjun Zhang, Yexiao Yun, Zhenfei Cao, Kai Mei, Zhenyu Chen 0001
IEEE Trans. Software Eng.1
2022 UniRLTest: universal platform-independent testing with reinforcement learning via image understanding
abstract
GUI testing has been prevailing in software testing. However, existing automated GUI testing tools mostly rely on frameworks of a specific platform. Testers have to fully understand platform features before developing platform-dependent GUI testing tools. Starting from the perspective of tester’s vision, we observe that GUIs on different platforms share commonalities of widget images and layout designs, which can be leveraged to achieve platform-independent testing. We propose UniRLTest, an automated software testing framework, to achieve platform independence testing. UniRLTest utilizes computer vision techniques to capture all the widgets in the screenshot and constructs a widget tree for each page. A set of all the executable actions in each tree will be generated accordingly. UniRLTest adopts a Deep Q-Network, a reinforcement learning (RL) method, to the exploration process and formalize the Android GUI testing problem to a Marcov Decision Process (MDP), where RL could work. We have conducted evaluation experiments on 25 applications from different platforms. The result shows that UniRLTest outperforms baselines in terms of efficiency and effectiveness.
Yulei Liu, Shengcheng Yu, Xin Li 0034, Yexiao Yun, Chunrong Fang, Zhenyu Chen 0001
ISSTA3
2022 SemCluster: a semi-supervised clustering tool for crowdsourced test reports with deep image understanding
abstract
Due to the openness of crowdsourced testing, mobile app crowdsourced testing has been subject to duplicate reports. The previous research methods extract the textual features of the crowdsourced test reports, combine with shallow image analysis, and perform unsupervised clustering on the crowdsourced test reports to clarify the duplication of crowdsourced test reports and solve the problem. However, these methods ignore the semantic connection between textual descriptions and screenshots, making the clustering results unsatisfactory and the deduplication effect less accurate.
Mingzhe Du, Shengcheng Yu, Chunrong Fang, Tongyu Li, Heyuan Zhang, Zhenyu Chen 0001
ESEC/SIGSOFT FSE2
2022 Test case prioritization using partial attention
Quanjun Zhang, Chunrong Fang, Weisong Sun, Shengcheng Yu, Yutao Xu, Yulei Liu
J. Syst. Softw.4
2021 Prioritize Crowdsourced Test Reports via Deep Screenshot Understanding
abstract
Crowdsourced testing is increasingly dominant in mobile application (app) testing, but it is a great burden for app developers to inspect the incredible number of test reports. Many researches have been proposed to deal with test reports based only on texts or additionally simple image features. However, in mobile app testing, texts contained in test reports are condensed and the information is inadequate. Many screenshots are included as complements that contain much richer information beyond texts. This trend motivates us to prioritize crowdsourced test reports based on a deep screenshot understanding. In this paper, we present a novel crowdsourced test report prioritization approach, namely DeepPrior. We fifirstrst represent the crowdsourced test reports with a novelly introduced feature, namely DeepFeature, that includes all the widgets along with their texts, coordinates, types, and even intents based on the deep analysis of the app screenshots, and the textual descriptions in the crowdsourced test reports. DeepFeature includes theBugFeature, which directly describes the bugs, and theContextFeature, which depicts the thorough context of the bug. The similarity of the DeepFeature is used to represent the test reports' similarity and prioritize the crowdsourced test reports. We formally define the similarity as DeepSimilarity. We also conduct an empirical experiment to evaluate the effectiveness of the proposed technique with a large dataset group. The results show that DeepPrior is promising, and it outperforms the state-of-the-art approach with less than half the overhead.
Shengcheng Yu, Chunrong Fang, Zhenfei Cao, Tongyu Li, Zhenyu Chen 0001
ICSE1
2021 Layout and Image Recognition Driving Cross-Platform Automated Mobile Testing
abstract
The fragmentation problem has extended from Android to different platforms, such as iOS, mobile web, and even mini-programs within some applications (app), like WeChat. In such a situation, recording and replaying test scripts is one of the most popular automated mobile app testing approaches. However, such approach encounters severe problems when crossing platforms. Different versions of the same app need to be developed to support different platforms relying on different platform supports. Therefore, mobile app developers need to develop and maintain test scripts for multiple platforms aimed at completely the same test requirements, greatly increasing testing costs. However, we discover that developers adopt highly similar user interface layouts for versions of the same app on different platforms. Such a phenomenon inspires us to replay test scripts from the perspective of similar UI layouts. In this paper, we propose an image-driven mobile app testing framework, utilizing Widget Feature Matching and Layout Characterization Matching to analyze app UIs. We use computer vision (CV) technologies to perform UI feature comparison and layout hierarchy extraction on mobile app screenshots to obtain UI structures containing rich contextual information of app widgets, including coordinates, relative relationship, etc. Based on acquired UI structures, we can form a platform-independent test script, and then locate the target widgets under test. Thus, the proposed framework non-intrusively replays test scripts according to a novel platform-independent test script model. We also design and implement a tool named LIRAT to devote the proposed framework into practice, based on which, we conduct an empirical study to evaluate the effectiveness and usability of the proposed testing framework. The results show that the overall replay accuracy reaches around 65.85% on Android (8.74% improvement over state-of-the-art approaches) and 35.26% on iOS (35% improvement over state-of-the-art approaches).
Shengcheng Yu, Chunrong Fang, Yexiao Yun, Yang Feng 0003
ICSE1
2020 STIFA: Crowdsourced Mobile Testing Report Selection Based on Text and Image Fusion Analysis
abstract
Crowdsourced mobile testing has been widely used due to its convenience and high efficiency [10]. Crowdsourced workers complete testing tasks and record results in test reports. However, the problem of duplicate reports has prevented the efficiency of crowdsourced mobile testing from further improving. Existing crowdsourced testing report analysis techniques usually leverage screenshots and text descriptions independently, but fail to recognize the link between these two types of information. In this paper, we present a crowdsourced mobile testing report selection tool, namely STIFA, to extract image and text feature information in reports and establish an image-text-fusion bug context. Based on text and image fusion analysis results, STIFA performs cluster analysis and report selection. To evaluate, we employed STIFA to analyze 150 reports from 2 apps. The results show that STIFA can extract, on average, 95.23% text feature information and 84.15% image feature information. Besides, STIFA reaches an accuracy of 87.64% in detecting duplicate reports. The demo can be found at https://youtu.be/Gw6ptqyQbQY.
Zhenfei Cao, Shengcheng Yu, Yexiao Yun, Chunrong Fang
ASE3
2019 From Data Quality to Model Quality: An Exploratory Study on Deep Learning
abstract
In the field of deep learning, people strive to construct high-quality deep neural networks (DNNs) to improve the accuracy of predicting. As well known, the quality of training data have great impacts on the quality of DNN models, since all the DNN models are obtained by training using these training data. However, there is not any reported systematic study on how the quality of training data affects the quality of DNN model. To study the relationships between data quality and model quality, we mainly consider four aspects of data quality including Skewed Classes, Sample Complexity, Label Quality, and Noisy Data in this paper. We design experiments on MNIST and Cifar-10, and attempt to find out the influences of four aspects on the quality of DNN models. Pearson correlation coefficient and Spearman correlation coefficient are utilized to evaluate such influences. Experimental results show that all the four aspects of data quality have significant impacts on the quality of DNN models. It means that the decrease of data quality in these four aspects will reduce the accuracy of the DNN models.
Tianxing He, Shengcheng Yu, Ziyuan Wang 0001, Jieqiong Li, Zhenyu Chen 0001
Internetware2
2019 Crowdsourced Report Generation via Bug Screenshot Understanding
abstract
Quality control is a challenge of crowdsourcing, especially in software testing. As some unprofessional workers involved, low-quality yieldings may hinder crowdsourced testing from satisfying requesters' requirements. Therefore, it is in demand to assist crowdworkers to raise bug report quality. In this paper, we propose a novel auxiliary method, namely CroReG, to generate crowdsourcing bug reports by analyzing bug screenshots uploaded by crowdworkers with image understanding techniques. The preliminary experiment results show that CroReG can effectively generate bug reports containing accurate screenshot captions and providing positive guidance for crowdworkers.
Shengcheng Yu
ASE1
2019 LIRAT: Layout and Image Recognition Driving Automated Mobile Testing of Cross-Platform
abstract
The fragmentation issue spreads over multiple mobile platforms such as Android, iOS, mobile web, and WeChat, which hinders test scripts from running across platforms. To reduce the cost of adapting scripts for various platforms, some existing tools apply conventional computer vision techniques to replay the same script on multiple platforms. However, because these solutions can hardly identify dynamic or similar widgets. It becomes difficult for engineers to apply them in practice. In this paper, we present an image-driven tool, namely LIRAT, to record and replay test scripts cross platforms, solving the problem of test script cross-platform replay for the first time. LIRAT records screenshots and layouts of the widgets, and leverages image understanding techniques to locate them in the replay process. Based on accurate widget localization, LIRAT supports replaying test scripts across devices and platforms. We employed LIRAT to replay 25 scripts from 5 application across 8 Android devices and 2 iOS devices. The results show that LIRAT can replay 88% scripts on Android platforms and 60% on iOS platforms. The demo can be found at: https: //github.com/YSC9848/LIRAT.
Shengcheng Yu, Chunrong Fang, Yang Feng 0003, Wenyuan Zhao, Zhenyu Chen 0001
ASE1
2019 An Exploratory Study on Judicial Image Quality Assessment Based on Deep Learning
abstract
Images are important judicial materials. With the deepening of intelligent systems in the judicial area, image quality plays a vital role in the result of many judicial applications. This paper firstly introduces deep learning into judicial image quality assessment. Pre-trained convolutional neural network (CNN) models are fine-tuned and then used to extract image features. Based on the features extracted from CNN models, we convert them into specific numbers representing the quality. A preliminary experiment has been designed and conducted on three types of judicial images. The experimental results show that our approach can outperform the existing image processing technique. Images used as investigation materials are more distinctive than the other two types, and they need an independent model for analyzing.
Weilin Cai, Shengcheng Yu, Zhenyu Chen 0001
QRS3