Jia Liu 0015

dblp:49/1245-15 · DBLP profile ↗
← Back
26ranked-venue papers
0as first author
15since 2021 · last 2026
0000-0001-8368-4898ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 15 · 8 since 2021Artificial intelligence and machine learning · 8 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 GUI test migration via LLM with scenario-granularity understanding
Shengcheng Yu, Chunrong Fang, Junyang Xing, Jia Liu 0015, Zhenyu Chen 0001
Frontiers Comput. Sci.6
2026 Test Script Intention Generation for Mobile Application via GUI Image and Code Understanding
abstract
Testing is the most direct and effective technique to ensure software quality. Test scripts always play a more important role in mobile app testing than test cases for source code, due to the GUI-intensive and event-driven characteristics of mobile applications (app). Test scripts focus on user interactions and the corresponding response events, which is significant for testing the target app functionalities. Therefore, it is critical to understand the test scripts for better script maintenance and modification. There exist some mature code understanding (i.e., code comment generation, code summarization) technologies that can be directly applied to functionality source code with business logic. However, such technologies will have difficulties when being applied to test scripts, because test scripts are loosely linked to Apps under Test (AUT) by widget selectors, and do not contain business logic themselves. In order to solve the test script understanding gap, this article presents a novel approach, namely TestIntention , to infer the intention of GUI test scripts. Test intention refers to the user expectations of app behaviors for specific operations . TestIntention formalizes test scripts with an operation sequence model. For each operation within the sequence, TestIntention extracts the target widget selector and links the selector to the GUI layout information or the corresponding response events. For widgets identified by XPath , TestIntention utilizes the image understanding technologies to explore the detailed information of the widget images, the intention of which is understood with a deep learning model. For widgets identified by ID , TestIntention first maps the selectors to the response methods with business logic, and then adopts code understanding technologies to describe code in natural language form. Results of all operations are combined to generate test intention for test scripts. An empirical experiment including different metrics proves the outstanding performance of TestIntention , outperforming baselines by much. Also, it is shown that TestIntention can save about 80% developers’ time to understand test scripts.
Shengcheng Yu, Chunrong Fang, Jia Liu 0015, Zhenyu Chen 0001
ACM Trans. Softw. Eng. Methodol.3
2026 Human Cognitive Pattern Simulation for Crowdsourced Test Report Consistency Detection
abstract
Crowdsourced testing has emerged as a prominent paradigm in software testing by leveraging the diversity of crowdworkers. In this paradigm, crowd-workers are required to submit a test report for each identified bug, which typically contains a textual description and a bug screenshot. However, due to varying worker expertise, many reports exhibit inconsistencies between the textual description and the bug screenshot, which hinder the report review process. Existing methods address this issue by automatically detecting report consistency, typically through matching the UI widgets referenced in the textual description with those visible in the bug screenshot. However, such methods focus only on surface-level element correspondence and fail to capture the abstract bug semantics, such as the functional meaning and bug-triggering context. Consequently, they lack the ability to detect more subtle but realistic inconsistencies. To bridge this gap, we propose INCONHUNTER, a novel method for crowdsourced test report consistency detection that explicitly simulates human cognitive pattern. In this pattern, humans typically adopt two complementary reasoning strategies. If the textual description allows them to form an expectation about the visual bug features, they assess consistency by verifying if the expected features appear in the bug screenshot. Otherwise, they shift to reasoning about if the bug-triggering context described in the report aligns with the app state shown in the bug screenshot. INCONHUNTERinstantiates this cognitive pattern through two LLM-powered modules, each dedicated to one reasoning strategy. We evaluate INCONHUNTERthrough experiments on our dataset with 2,310 labeled crowdsourced test reports, and results show that INCONHUNTERoutperforms baselines by 14.00%–19.28%, demonstrating superior effectiveness, monetary-based cost efficiency, and alignment with human cognitive pattern.
Yuchen Ling, Shengcheng Yu, Shuguang Chen, Liuming Wang, Chunrong Fang, Jia Liu 0015, Zhenyu Chen 0001
IEEE Trans. Software Eng.7
2025 Lightweight Probabilistic Coverage Metrics for Efficient Testing of Deep Neural Networks
abstract
Deep neural networks (DNNs) have been deployed in many software systems to assist in various tasks.Accompanying with great performance, however, DNNs could also exhibit erroneous behaviors and cause massive losses.To assist the quality assurance and measure the testing adequacy of DNNs, recent research has proposed many neuron coverage (NC) metrics that measure the proportion of neurons activated in executions.While neuron coverage metrics are an analogy to structural code coverage for conventional software programs and reflect the internal behaviors of DNN models in executions, we still lack a comprehensive understanding about the application effectiveness of neuron coverage for deep learning testing.Besides, technologies like DeepGini and ATS have demonstrated the superiority of output probability vectors over neuron coverage for test selection, these techniques do not serve as coverage metrics and thus cannot be directly compared with neuron coverage in other deep learning testing tasks.This paper systematically evaluates the effectiveness of neuron activation-based coverage in multiple testing application scenarios.In addition, to better understand neuron coverage bottlenecks, we further propose an output-probability vector-based coverage metric (named Pt) inspired by existing test selection technique.We perform a comprehensive experiments across three prevalent application scenarios: assessing dataset diversity, improving model retraining, and guiding test generation.Experimental results show that most neuron coverage techniques are not very effective in deep learning testing.Coverage based on neuron activation state do not improve testing efficiency like code coverage.In contrast, the output-based coverage we introduced demonstrates significantly enhanced effectiveness.Our study improves the comprehension of * Yang Feng is the corresponding author.
Yining Yin, Yang Feng 0003, Shihao Weng, Jia Liu 0015
Internetware5
2025 A Large-Scale Empirical Study of Actionable Warning Distribution Within Projects
abstract
Static Analysis Tools (SATs) show potential defect detection ability while their usability is severely hindered by massive unactionable warnings. To improve the usability of SATs, many machine learning-based Actionable Warning Identification (AWI) studies have been proposed, which mainly focus on mining warning features and improving identification models to identify actionable warnings. However, the underlying distribution of the warning dataset, which is closely related to feature mining and thereby affects AWI model performance, is not well-explored by these studies. Further, there is a lack of a well-prepared warning dataset to support the distribution analysis. In this article, we first propose a warning dataset construction approach, which incorporates manual inspection and verification latency into postprocess labels from an advanced closed-warning heuristic and thereby acquire credible labels. Based on 10 large-scale and real-world projects with 25K+ revisions and 2087K+ SpotBugs warnings, we construct a qualified warning dataset with 11975 distinct warnings. Subsequently, we thoroughly analyze the actionable warning distribution within projects against our constructed dataset from six warning characteristics (i.e., category, type, priority, rank, file, and method). Based on the analysis results, we present 16 findings. Finally, a preliminary study demonstrates that our findings can be practical and instructive in improving the usability of SATs.
Xiuting Ge, Chunrong Fang, Xuanye Li, Jia Liu 0015, Zhenyu Chen 0001
IEEE Trans. Dependable Secur. Comput.5
2025 NLPLego: Assembling Test Generation for Natural Language Processing Applications
abstract
With the development of Deep Learning, Natural Language Processing (NLP) applications have reached or even exceeded human-level capabilities in certain tasks. Although NLP applications have shown good performance, they can still have bugs like traditional software and even lead to serious consequences. Inspired by Lego blocks and syntax structure analysis, we propose an assembling test generation method for NLP applications or models and implement it in NLPLego . The key idea of NLPLego is to assemble the sentence skeleton and adjuncts in order by simulating the building of Lego blocks to generate multiple grammatically and semantically correct sentences based on one seed sentence. The sentences generated by NLPLego have derivation relations and different degrees of variation. These characteristics make it well-suited for integration with metamorphic testing theory, addressing the challenge of test oracle absence in NLP application testing. To validate NLPLego , we conduct experiments on three commonly used NLP tasks (i.e., machine reading comprehension, sentiment analysis, and semantic similarity measures), focusing on the efficiency of test generation and the quality and effectiveness of generated tests. We select five advanced NLP models and one popular industrial NLP software as the tested subjects. Given seed tests from SQuAD 2.0, SST, and QQP, NLPLego successfully detects 1,732, 3,140, and 261,879 incorrect behaviors with around 93.1% precision in three tasks, respectively. The experiment results show that NLPLego can efficiently generate high-quality tests for multiple NLP tasks to detect erroneous behaviors effectively. In the case study, we analyze the testing results provided by NLPLego to obtain intuitive representations of the different NLP capabilities of the tested subjects. The case study confirms that NLPLego can provide developers with clarity on the direction to improve NLP models or applications, laying the foundation for enhancing performance.
Pin Ji, Yang Feng 0003, Ruohao Zhang, Ruichen Xue, Weitao Huang, Jia Liu 0015
ACM Trans. Softw. Eng. Methodol.7
2025 MoCo: Fuzzing Deep Learning Libraries via Assembling Code
abstract
The rapidly developing Deep Learning (DL) techniques have been applied in software systems of various types. However, they can also pose new safety threats with potentially serious consequences, especially in safety-critical domains. DL libraries serve as the underlying foundation for DL systems, and bugs in them can have unpredictable impacts that directly affect the behaviors of DL systems. Previous research on fuzzing DL libraries still has limitations in generating tests corresponding to crucial testing scenarios and constructing test oracles. In this paper, we proposeMoCo, a novel fuzzing testing method for DL libraries via assembling code. The seed tests used byMoCoare code files that implement DL models, covering both model construction and training in the most common real-world application scenarios for DL libraries.MoCofirst disassembles the seed code files to extract templates and code blocks, then applies code block mutation operators (e.g., API replacement, random generation, and boundary checking) to generate new code blocks that fit the template. To ensure the correctness of the code block mutation, we employ the Large Language Model to parse the official documents of DL libraries for information about the parameters and the constraints between them. By inserting context-appropriate code blocks into the template,MoCocan generate a tree of code files with intergenerational relations. According to the derivation relations in this tree, we construct the test oracle based on the execution state consistency and the calculation result consistency. Since the granularity of code assembly is controlled rather than randomly divergent, we can quickly pinpoint the lines of code where the bugs are located and the corresponding triggering conditions. We conduct a comprehensive experiment to evaluate the efficiency and effectiveness ofMoCousing three widely-used DL libraries (i.e., TensorFlow, PyTorch, and Jittor). During the experiments,MoCodetects 77 new bugs of four types in three DL libraries, where 55 bugs have been confirmed, and 39 bugs have been fixed by developers. The experimental results demonstrate thatMoCocan generate high-quality tests that cover crucial testing scenarios and detect different types of bugs, which helps developers improve the reliability of DL libraries.
Pin Ji, Yang Feng 0003, Duo Wu, Lingyue Yan, Penglin Chen, Jia Liu 0015
IEEE Trans. Software Eng.6
2024 Datactive: Data Fault Localization for Object Detection Systems
abstract
Object detection (OD) models are seamlessly integrated into numerous intelligent software systems, playing a crucial role in various tasks. These models are typically constructed upon humanannotated datasets, whose quality can greatly affect their performance and reliability. Erroneous and inadequate annotated datasets can induce classification/localization inaccuracies during deployment, precipitating security breaches or traffic accidents that inflict property damage or even loss of life. Therefore, ensuring and improving data quality is a crucial issue for the reliability of the object detection system. This paper introduces Datactive, a data fault localization technique for object detection systems. Datactive is designed to locate various types of data faults including mislocalization and missing objects, without utilizing the prediction of object detection models trained on dirty datasets. To achieve this, we first construct foreground-only and background-included datasets via data disassembling strategies, and then employ a robust learning method to train classifiers using disassembled datasets. Based on the classifier predictions, Datactive produces a unified suspiciousness score for both foreground annotations and image backgrounds. It allows testers to easily identify and correct faulty or missing annotations with minimal effort. To validate the effectiveness, we conducted experiments on three datasets with 6 baselines, and demonstrated the superiority of Datactive from various aspects. We also explored Datactive's ability to find natural data faults and its application in both training and evaluation scenarios.
Yining Yin, Yang Feng 0003, Shihao Weng, Yuan Yao 0001, Jia Liu 0015
ISSTA5
2024 Improving actionable warning identification via the refined warning-inducing context representation
Xiuting Ge, Chunrong Fang, Xuanye Li, Quanjun Zhang, Jia Liu 0015, Zhenyu Chen 0001
Sci. China Inf. Sci.5
2024 Benchmarking Object Detection Robustness against Real-World Corruptions
Zhijie Wang 0014, Lei Ma 0003, Chunrong Fang, Tongtong Bai, Xufan Zhang, Jia Liu 0015, Zhenyu Chen 0001
Int. J. Comput. Vis.7
2023 Prioritizing Testing Instances to Enhance the Robustness of Object Detection Systems
abstract
Object detection models have been widely deployed in military and life-related intelligent software systems. However, along with the outstanding success of object detection, it may exhibit abnormal behavior and lead to severe accidents and losses. During the development and evaluation process, training and evaluating an object detection model are computationally intensive, while preparing annotated tests requires extremely heavy manual labor. Therefore, reducing the annotation budget of test data collection becomes a challenging and necessary task. Although many test prioritization approaches for DNN-based systems have been proposed, the large differences between classification and object detection make them difficult to apply to testing object detection models.
Shihao Weng, Yang Feng 0003, Yining Yin, Jia Liu 0015
Internetware4
2023 An unsupervised feature selection approach for actionable warning identification
Xiuting Ge, Chunrong Fang, Jia Liu 0015, Mingshuang Qing, Xuanye Li
Expert Syst. Appl.3
2023 An Empirical Study of Class Rebalancing Methods for Actionable Warning Identification
abstract
Actionable warning identification (AWI) is crucial for improving the usability of static analysis tools. Currently, machine learning (ML)-based AWI approaches are notably common, which mainly focus on seeking high performance by improving the warning feature extraction and advancing the AWI model training. However, these approaches ignore an important fact that the number of actionable warnings is much smaller than that of unactionable warnings in the warning dataset used for the AWI model training (i.e., the class imbalance). Learning from such an imbalanced dataset may limit the performance of ML-based AWI approaches. To bridge the above gap, we are the first to conduct a comprehensive empirical study to investigate the impact of class imbalance on the ML-based AWI performance, whether class rebalancing methods can improve the ML-based AWI performance, and the differences of class rebalancing methods in the ML-based AWI model. Our empirical study is performed on 9 real-world and large-scale warning datasets, 25 typical class rebalancing methods, and 7 commonly used ML models. The experimental results show that 1) the class imbalance has a negative impact on the ML-based AWI performance; 2) 85% class rebalancing methods can significantly improve the ML-based AWI performance, but 8% ones do not work in the imbalanced warning datasets; 3) RandomOverSampler combined with AdaBoost/Random Forest can make the ML-based AWI model achieve optimal performance on nine warning datasets. Finally, we provide three practical guidelines that could help refine ML-based AWI approaches.
Xiuting Ge, Chunrong Fang, Tongtong Bai, Jia Liu 0015
IEEE Trans. Reliab.4
2021 Predoo: precision testing of deep learning operators
abstract
Deep learning(DL) techniques attract people from various fields with superior performance in making progressive breakthroughs. To ensure the quality of DL techniques, researchers have been working on testing and verification approaches. Some recent studies reveal that the underlying DL operators could cause defects inside a DL model. DL operators work as fundamental components in DL libraries. Library developers still work on practical approaches to ensure the quality of operators they provide. However, the variety of DL operators and the implementation complexity make it challenging to evaluate their quality. Operator testing with limited test cases may fail to reveal hidden defects inside the implementation. Besides, the existing model-to-library testing approach requires extra labor and time cost to identify and locate errors, i.e., developers can only react to the exposed defects. This paper proposes a fuzzing-based operator-level precision testing approach to estimate individual DL operators' precision errors to bridge this gap. Unlike conventional fuzzing techniques, valid shape variable inputs and fine-grained precision error evaluation are implemented. The testing of DL operators is treated as a searching problem to maximize output precision errors. We implement our approach in a tool named Predoo and conduct an experiment on seven DL operators from TensorFlow. The experiment result shows that Predoo can trigger larger precision errors compared to the error threshold declared in the testing scripts from the TensorFlow repository.
Xufan Zhang, Chunrong Fang, Jia Liu 0015, Dong Chai, Zhenyu Chen 0001
ISSTA5
2021 Duo: Differential Fuzzing for Deep Learning Operators
abstract
Deep learning (DL) libraries reduce the barriers to the DL model construction. In DL libraries, various building blocks are DL operators with different functionality, responsible for processing high-dimensional tensors during training and inference. Thus, the quality of operators could directly impact the quality of models. However, existing DL testing techniques mainly focus on robustness testing of trained neural network models and cannot locate DL operators’ defects. The insufficient test input and undetermined test output in operator testing have become challenging for DL library developers. In this article, we propose an approach, namely Duo, which combines fuzzing techniques and differential testing techniques to generate input and evaluate corresponding output. It implements mutation-based fuzzing to produce tensor inputs by employing nine mutation operators derived from genetic algorithms and differential testing to evaluate outputs’ correctness from multiple operator instances. Duo is implemented in a tool and used to evaluate seven operators from TensorFlow, PyTorch, MNN, and MXNet in an experiment. The result shows that Duo can expose defects of DL operators and realize multidimension evaluation for DL operators from different DL libraries.
Xufan Zhang, Chunrong Fang, Jia Liu 0015, Dong Chai, Zhenyu Chen 0001
IEEE Trans. Reliab.5
2017 An empirical study on user-topic rating based collaborative filtering methods
Tieke He, Zhenyu Chen 0001, Jia Liu 0015, Xiaofang Zhou 0001, Xingzhong Du, Weiqing Wang 0001
World Wide Web3
2016 Mining Feature-Opinion from Reviews Based on Dependency Parsing
abstract
Manually reading all the product reviews to find a satisfying item is not only labor-intensive, but also tedious for the consumers.In this paper, we propose a feature-opinion mining approach to automatically summarize the reviews.Specifically, in our approach we first utilize a regression model to generate sentiment word, including phrase and its sentiment weight, then extract feature based on the dependency relationship between feature word and sentiment word, and finally we assign score to feature according to the dependency relationship.The experimental results demonstrate that our approach can effectively mine the featureopinion from reviews.
Tieke He, Jia Liu 0015
SEKE4
2016 Exploring the Influence of Time Factor in Bug Report Prioritization
abstract
Time factor has been widely applied into a wide range of data mining areas, such as social network and information retrieval.The main idea of taking time factor into consideration is that human activities may have some relations to time pattern.However, little attention has been pulled on the study of time factor in the area of software engineering.In this paper, we endeavour to explore to what extent time factor affects the prioritization of bug reports, a specified while important task in software engineering.Specifically, we test four time factors that may have some influence on this task, which are time of day, normal time, day of week, and days to major version.After the validation of relatedness of all these factors, we conduct an extensive set of experiments on two datasets to verify the effectiveness of these factors.The experimental results demonstrate that we can effectively improve the results by metrics of both Precision and Recall, on two classical models, i.e., the SVM model and Naive Bayes model.
Zhengjie Xu, Tieke He, Jia Liu 0015, Zhenyu Chen 0001
SEKE5
2016 Mubug: a mobile service for rapid bug tracking
Yang Feng 0003, Mengyu Dou, Jia Liu 0015, Zhenyu Chen 0001
Sci. China Inf. Sci.4
2016 Mining Feature-Opinion from Reviews Based on Dependency Parsing
abstract
The manual reading of all the product reviews to find a satisfying item is not only labor-intensive, but also tedious for the consumers. In this paper, we propose a feature-opinion mining approach to automatically summarize the reviews, which is based on dependency parsing. Specifically, in our approach we first utilize a regression model to generate sentiment word, including phrase and its sentiment weight, and then we extract the feature based on the dependency relationship between feature word and sentiment word, finally we assign a score to the feature according to the dependency relationship. The experimental results demonstrate that our approach can effectively mine the feature-opinion from reviews.
Tieke He, Jia Liu 0015
Int. J. Softw. Eng. Knowl. Eng.4
2014 Testing as an Investment
Chunrong Fang, Jia Liu 0015, Zhenyu Chen 0001
SEKE4
2014 Developer social networks in software engineering: construction, analysis, and applications
Liming Nie, He Jiang 0001, Zhenyu Chen 0001, Jia Liu 0015
Sci. China Inf. Sci.5
2013 Motivating and orienting novice students to value introductory software engineering
abstract
Students with little professional software development experience typically have low intrinsic motivation and beyond achieving a good grade, low extrinsic motivation to study and appreciate the value of software engineering curricula. Unlike other subjects, introductory software engineering instructors exert a great deal of effort justifying and motivating their course topics. Since 2006 the University of Hawaii has been developing a series of “early awareness” engagements within an introductory Systems Analysis and Design course designed to foster intrinsic and extrinsic motivation and orient students to value learning software engineering. Measuring motivation and perceived value is difficult, but there are key indicators to determine the impact of our improvements such as higher course evaluations, greater class attendance, and increased positive feedback from students and employers. Our results show that these key indicators have improved since introducing early engagements. Furthermore, students like this approach, value the course more, do better quality work, and evaluate the course more positively. Longitudinal follow-ups indicate greater interest and success in pursuing software engineering related careers. This paper shares the details of our early awareness engagements, how they are applied in the classroom, some of our experiences in using them, and evidence that they have a positive impact. Our goal is to provide specific and practical means that instructors can use immediately to improve the perceived value of software engineering for novice students prior to or while they study it.
Daniel Port, Chris Rachal, Jia Liu 0015
CSEE&T3
2013 ABEY: an Incremental Personalized Method Based on Attribute Entropy for Recommender Systems (S)
Xingzhong Du, Tieke He, Zhenyu Chen 0001, Jia Liu 0015, Chengfeng Hui
SEKE4
2013 Comparing Collaborative Filtering Methods Based on User-Topic Ratings
Tieke He, Xingzhong Du, Weiqing Wang 0001, Zhenyu Chen 0001, Jia Liu 0015
SEKE5
2012 An Empirical Study on Recommendation Methods for Vertical B2C E-commerce
Chengfeng Hui, Jia Liu 0015, Zhenyu Chen 0001, Xingzhong Du, Weiyun Ma
SEKE2