VLDB 2026 Research / reviewers in the wild / expert
Junxia Guo
dblp:38/8024
· DBLP profile ↗
25ranked-venue papers
9as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 19 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-authorArtificial intelligence and machine learning · 4 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REE: A Cooperative Framework for Robust and Explainable Vulnerability Detection
Lele Zhong, Fa Zhong, Junxia Guo |
COMPSAC | 3 |
| 2025 | Intent Based E2E Automated Test Case Generation for Web Applications Using LLMabstractEnsuring reliability in dynamic web applications is like hitting a moving target, where content constantly changes in response to user interactions and real-time updates. Therefore, testing these applications automatically remains challenging. Additionally,elements, which render graphical content outside the DOM and lack semantic metadata, makes it difficult to automatically extract and validate their content, leading to incomplete testing in existing methods. This paper proposes a novel approach named E2E-TestGen, which based on intent-driven testing, leverages Large Language Models (LLMs). E2ETestGen interacts dynamically with the DOM by triggering JavaScript-driven updates and tracking real-time state transitions, ensuring comprehensive validation across all dynamic elements and interactions to improve coverage. For canvas-rendered graphical components, we integrate an image-based object detection and text extraction module to extract meaningful data from visual content, enabling the generation of realistic test case scenarios guided by user intent. The experiments on multiple open-source web applications are conducted, which show E2E-TestGen achieving 88.5% and 80.1% code and branch coverage, respectively, along with a 27.91% improvement in fault detection over baseline approaches. Zain Ul Abideen 0008, Junxia Guo |
COMPSAC | 2 |
| 2025 | Peft: Multiline Complex Patch Correctness Assessment Based on Fine-Tuning Large Language Model with "Golden Data"abstractIn recent years, automated program repair has garnered considerable attention owing to its potential to mitigate software maintenance costs. There remain challenges for automated program repair techniques to be widely applied in practice, where too many overfitting patches are one of the issues. Automated patch correctness assessment emerges as a crucial means to enhance the viability of recommended patches in automated program repair technology. The current automated patch correctness assessment methods exhibit suboptimal performance when evaluating multiline patches compared to their effectiveness in assessing single-line patches. This disparity arises because multiline patches involve intricate contextual semantic relationships and significant program structure modifications, such as branching and looping. Those complexities render it more challenging for neural networks to capture the latent features essential for assessing correctness. To address the above problem, this paper proposes an automated patch correctness assessment method, named Peft, which fine-tunes a large language model using “golden data” that is generated based on semantic and program structural features, resulting in a patch evaluation model. Experiments were conducted on three datasets with varying proportions of multiline complex patches. The patches were sourced from real-world applications or generated by 23 APR techniques for fixing Defects4J v1.2. The results demonstrated that Peft consistently outperformed existing methods across all datasets. Notably, when evaluating datasets composed entirely of multiline complex patches, Peft significantly outperformed the open-source baseline technique Cache (accuracy: 79.2% vs. 66.1%, F1-score: 81.9% vs. 69.5%). Xiaoxi Zheng, Ruilian Zhao, Junxia Guo |
QRS | 3 |
| 2024 | Multi-Objective Test Case Generation for Web Applications with Limited ResourcesabstractWith the development of web technology, web applications are becoming more complicated and diverse, which makes full testing need lots of resources. Therefore, industries often face how to do effective web application testing with limited time and resources. Evaluating the importance of a web application's page nodes and testing high-importance nodes first is a workable solution. However, the existing methods only consider the content importance on the view of static and local, which lack the consideration of dynamic execution process and the overall topology structure. Web applications usually contain a large number of interaction events with users, which need to dynamically execute server-side logic code to respond to them and update the page nodes. Meanwhile, whether the nodes are in the key position of information transmission in the topology structure of the whole model is also very important for evaluating the importance of web nodes. Thus, in this paper, we propose a novel evaluation method of node comprehensive importance for web applications, which introduces the server-side code coverage to reflect the dynamic information of web applications, the betweenness centrality to consider topology information from a global perspective, and uses the multi-attribute decision-making method to calculate the final comprehensive importance. In addition, this paper implements multi-objective test case generation with GA and NSG A2 algorithms guided by diversity and page importance under the condition of limited resources. This can ensure that the generated test suite has good quality, while also the nodes with higher importance degrees are tested first. In detail, it can cover all important nodes using fewer test cases, with an average reduction of 1.14, compared with the existing methods. Junxia Guo |
COMPSAC | 1 |
| 2024 | EFSM Model-Based Testing for Android ApplicationsabstractModel-based testing provides an effective means for ensuring the quality of Android apps. Nevertheless, existing models that focus on event sequences and abstract them into Finite State Machines (FSMs) may lack precision and determinism because of the different data values of events that can result in various states of Android applications. To address this issue, a novel model based on Extended Finite State Machines (EFSMs) for Android apps is proposed in this paper. The approach leverages machine learning to infer data constraints on events and annotates them on state transitions, leading to a more precise and deterministic model. Additionally, a state abstraction strategy is presented to further refine the model. Besides, test diversity plays a vital role in enhancing test suite effectiveness. To achieve high coverage and fault detection, test cases are generated from the EFSM model with the help of a Genetic Algorithm (GA), guided by test diversity. To evaluate the effectiveness of our approach, this paper carries out experiments on 93 open-source apps. The results show that our approach performs better in code coverage and crash detection than the existing open-source model-based testing tools. Particularly, the 19 unique crashes that involve complex data constraints are detected by our approach. Junxia Guo, Beite Li, Ruilian Zhao |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2023 | Research on hyper-level of hyper-heuristic framework for MOTCPabstractSummary Heuristic algorithms are widely used to solve multi‐objective test case prioritization (MOTCP) problems. However, they perform differently for different test scenarios, which conducts difficulty in applying a suitable algorithm for new test requests in the industry. A concrete hyper‐heuristic framework for MOTCP (HH‐MOTCP) is proposed for addressing this problem. It mainly has two parts: low‐level encapsulating various algorithms and hyper‐level including an evaluation and selection mechanism that dynamically selects low‐level algorithms. This framework performs good but still difficult to keep in the best three. If the evaluation mechanism can more accurately analyse the current results, it will help the selection strategy to find more conducive algorithms for evolution. Meanwhile, if the selection strategy can find a more suitable algorithm for the next generation, the performance of the HH‐MOTCP framework will be better. In this paper, we first propose new strategies for evaluating the current generation results, then perform an extensive study on the selection strategies which decide the heuristic algorithm for the next generation. Experimental results show that the new evaluation and selection strategies proposed in this paper can make the HH‐MOTCP framework more effective and efficient, which makes it almost the best two except for one test object and ahead in about 40% of all test objects. Junxia Guo, Jinjin Han, Zheng Li 0002 |
Softw. Test. Verification Reliab. | 1 |
| 2020 | WLTDroid: Repackaging Detection Approach for Android Applications
Junxia Guo, Rilian Zhao, Zheng Li 0002 |
WISA | 1 |
| 2020 | Improving Detection Accuracy for Malicious JavaScript Using GAN
Junxia Guo, Qiyun Cao, Rilian Zhao, Zheng Li 0002 |
ICWE | 1 |
| 2020 | Automatic Model Completion for Web Applications
Ruilian Zhao, Chen Chen 0064, Junxia Guo |
ICWE | 4 |
| 2020 | Dynamic Time Window based Reward for Reinforcement Learning in Continuous Integration TestingabstractContinuous Integration (CI) testing is an expensive, time-consuming, and resource-intensive process. Test case prioritization (TCP) can effectively reduce the workload of regression testing in the CI environment, where Reinforcement Learning (RL) is adopted to prioritize test cases, since the TCP in CI testing can be formulated as a sequential decision-making problem, which can be solved by RL effectively. A useful reward function is a crucial component in the construction of the CI system and a critical factor in determining RL’s learning performance in CI testing. This paper focused on the validity of the execution history information of the test cases on the TCP performance in the existing CI testing optimization methods based on RL, and a Dynamic Time Window based reward function are proposed by using partial information dynamically for fast feedback and cost reduction. Experimental studies are carried out on six industrial datasets. The experimental results showed that using dynamic time window based reward function can significantly improve the learning efficiency of RL and the fault detection ability when comparing with the reward function based on fixed time window. Chaoyue Pan, Yang Yang 0099, Zheng Li 0002, Junxia Guo |
Internetware | 4 |
| 2020 | Convergence based Evaluation Strategies for Learning Agent of Hyper-heuristic Framework for Test Case PrioritizationabstractLearning agent plays significant role in the hyper-heuristic framework for test case prioritization, where an evaluation strategy is applied to evaluate the execution results produced by the current heuristic algorithm and select the most appropriate heuristic algorithm for the next generation. Hierarchical Distribution (HD) is used as evaluation strategy based on the dominance relationship between the individuals from the present and last generations. In addition to the distribution of the solution set, a good convergence towards the optimal Pareto front is often desired. In this paper, the convergence ability of the individuals is further considered in the design of the evaluation strategy for the learning agent, in which Pareto Dominance and Convergence Information are adopted. Three evaluation strategies are proposed and empirically studied, and the experimental results show that the hyper-heuristic algorithms with the proposed evaluation strategies are more effective and efficient for test case prioritization. Jinjin Han, Zheng Li 0002, Junxia Guo, Ruilian Zhao |
QRS | 3 |
| 2020 | Thread Scheduling Sequence Generation Based on All Synchronization Pair Coverage CriteriaabstractTesting multi-thread programs becomes extremely difficult because thread interleavings are uncertain, which may cause a program getting different results in each execution. Thus, Thread Scheduling Sequence (TSS) is a crucial factor in multi-thread program testing. A good TSS can obtain better testing efficiency and save the testing cost especially with the increase of thread numbers. Focusing on the above problem, in this paper, we discuss a kind of approach that can efficiently generate TSS based on the concurrent coverage criteria. First, we give a definition of Synchronization Pair (SP) as well as all Synchronization Pairs Coverage (ASPC) criterion. Then, we introduce the Synchronization Pair Thread Graph (SPTG) to describe the relationships between SPs and threads. Moreover, this paper presents a TSS generation method based on the ASPC according to SPTG. Finally, TSSs automatic generation experiments are conducted on six multi-thread programs in Java Library with the help of Java Path Finder (JPF) tool. The experimental results illustrate that our method not only generates TSSs to cover all SPs but also requires less state number, transition number as well as TSS number when satisfying ASPC, compared with other three widely used TSS generation methods. As a result, it is clear that the efficiency of TSS generation is obviously improved. Junxia Guo, Zheng Li 0002, CunFeng Shi, Ruilian Zhao |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2019 | Class Imbalance Data-Generation for Software Defect PredictionabstractThe imbalanced nature of class in software defect data, which including intra-class imbalance and inter-classes imbalance, increases the difficulty of learning an effective defect prediction model. Most of sampling and example generation approaches just focused on inter-class imbalanced defect data, and they are not effective to handle the issue of intra-class imbalance. This paper proposed a distribution based data generation approach for software defect prediction to deal with inter-class and intra-class imbalanced data simultaneously. First, the classified sub-regions are clustered according to the distribution in the sample feature space. Second, the data are generated by corresponding strategies according to different distribution in sub-regions, where the inter-class balance is achieved by increasing the number of defective samples, and the intra-class balance is achieved by generating different density of data in different sub-regions. Experiment results show that the proposed method can reduce the impact of data imbalance on defect prediction and improve the accuracy of software defect prediction model effectively by generating inter-class and intra-class balanced defects data. Zheng Li 0002, Junxia Guo |
APSEC | 3 |
| 2019 | Research on Page Object Generation Approach for Web Application TestingabstractTest code generated by using the page object design pattern during web testing is easy to maintain.Page clustering is an essential stage of the page object approach.However, existing methods only consider the DOM structure in page clustering, which leads to inaccuracy when generating page objects.A state with the same DOM structure may result in an entirely different migration.The method of considering only the DOM structure cannot accurately generate page object classes.In order to improve the accuracy of page object generation, this paper not only considers DOM structure information but also considers CSS styles and the attributes of DOM elements in page clustering.Based on the experimental evaluation results, our method can automatically generate page objects that cover most of the application functions, which is more effective for the creation and maintenance of web test cases. Yimei Chen, Zheng Li 0002, Ruilian Zhao, Junxia Guo |
SEKE | 4 |
| 2018 | Search-Based Efficient Automated Program Repair Using Mutation and Fault LocalizationabstractProgram faults are unavoidable phenomena in the software development. The application of mutation and fault localization techniques is effective in automated program repair, but suffers the inevitable high execution cost as the number of mutants increases exponentially for large industrial programs. The combination with fault localization techniques can reduce the cost by mutating the statements with high suspicious values first. However, the accuracy of fault localization techniques has not been good enough for real applications, and some faults may occur on a statement which is related to the statement with a high suspicious value rather than itself. Therefore, the greedy strategy used currently may not be effective, resulting in inefficient repairs. Finding a mutant as a correct patch should be regarded as a continuous process of searching the global solutions. In this paper, we proposed the search-based automated program repair using mutation and fault localization, which not only takes advantage of fault localization but also overcomes the disadvantage of the greedy strategy used in the mutation generation. The initial population of the search algorithm is constructed by the mutants generated from the statements with high suspicious values using the fault localization, which is a set of rough solutions. A hybrid-crossover operator is then designed where the fixed position crossover operator is used to converge to the global optimal solutions and the random position crossover operator is used to explore the entire search space faster, respectively. The experimental results on the Siemens suite indicate that the proposed approach can improve the efficiency with the same effectiveness compared to the exhaustive approach, and show that the non-random initial population method and the hybrid-crossover strategy can improve the efficiency of the search process. Shuyao Sun, Junxia Guo, Ruilian Zhao, Zheng Li 0002 |
COMPSAC (1) | 2 |
| 2018 | EFSM-Oriented Minimal Traces Set Generation Approach for Web ApplicationsabstractMost of web applications models focus on sequencing of events, where the ignored parameters or DOM elements changes and the relationship between the execution conditions and web states are crucial for analyzing and testing the behavior of client-side of client-server web applications. In this paper, we first define a novel trace, which can represent dynamic behaviors of web applications more accurately. Then an EFSM-oriented minimal traces set generation approach is proposed for modelling web applications. In order to ensure the integrity of the EFSM model and improve the effectiveness of the modelling process, three adequacy criteria with respect to all events, JS branches and DOM structures, are applied to compensate the traces and to guide the minimal traces set generation by greedy algorithm. Finally, the minimal traces set is abstracted into an EFSM as the behavior model for web applications. We implement a prototype tool for the proposed approach and empirically evaluate that the minimal traces set generation approach based on two web applications. The results show that the traces generated by the approach is effective and the all JavaScript branches coverage criteria is most appropriate to select the minimal traces set used for modelling. Junxia Guo, Zheng Li 0002, Ruilian Zhao |
COMPSAC (1) | 2 |
| 2018 | A Test Case Generation Method Based on State Importance of EFSM for Web ApplicationabstractTest cases generation is a principal process in web application testing.Most existing methods generate test cases for improving test efficiency mainly from the aspects like minimizing the test case suite, increasing the code coverage, and so on.However, similar with traditional software having important functions, classes or modules, some web states are more vital than others in web applications.It can be thought that those vital web states relatively have higher influence on the performance of web application.So, they should be given more attention in test case generation.In more detail, the importance of web states can be measured from its page contents or topological structures.Meanwhile, as we known Model-based Testing is a kind of widely used approach in automatic test case generation.Therefore in this paper we propose an EFSM based test case generation method considering the importance of web states for web applications.The experimental results show that our methods can deterministically enhance the testing efficiency of web application. Junxia Guo, Linjie Sun, Zheng Li 0002, Ruilian Zhao |
SEKE | 1 |
| 2018 | Concrete hyperheuristic framework for test case prioritizationabstractAbstract Test case prioritization (TCP), which aims to find the optimal test case execution sequences for specific testing objects, has been widely used in regression testing. A wide variety of search methodologies and algorithms have been proposed to optimize test case execution sequences, namely, search‐based TCP. However, different algorithms perform differently and have different implementation costs and specific situations where an algorithm usually performs with high effectiveness and efficiency. When facing a new testing scenario, it is actually difficult to decide which algorithm is suitable. In this paper, to address the algorithm selection problem for different test scenarios, a more generally applicable algorithm based on a hyperheuristic strategy is proposed for search‐based TCP. This includes a range of multiobjective algorithms with a variety of crossover strategies and a learning agent strategy to evaluate and select the appropriate algorithm execution sequence dynamically for different scenarios. The concrete hyperheuristic framework for multiobjective TCP is presented with an algorithm's repository in the low level and the learning agent strategy in the higher level. Experiments show that the proposed learning agent strategy can accurately evaluate algorithms in multiobjective problems and select the appropriate algorithm in each iteration. Zheng Li 0002, Junxia Guo, Ruilian Zhao |
J. Softw. Evol. Process. | 3 |
| 2017 | Smooth Neuroadaptive PI Tracking Control of Nonlinear Systems With Unknown and Nonsmooth Actuation CharacteristicsabstractThis paper considers the tracking control problem for a class of multi-input multi-output nonlinear systems subject to unknown actuation characteristics and external disturbances. Neuroadaptive proportional-integral (PI) control with self-tuning gains is proposed, which is structurally simple and computationally inexpensive. Different from traditional PI control, the proposed one is able to online adjust its PI gains using stability-guaranteed analytic algorithms without involving manual tuning or trial and error process. It is shown that the proposed neuroadaptive PI control is continuous and smooth everywhere and ensures the uniformly ultimately boundedness of all the signals of the closed-loop system. Furthermore, the crucial compact set precondition for a neural network (NN) to function properly is guaranteed with the barrier Lyapunov function, allowing the NN unit to play its learning/approximating role during the entire system operation. The salient feature also lies in its low complexity in computation and effectiveness in dealing with modeling uncertainties and nonlinearities. Both square and nonsquare nonlinear systems are addressed. The benefits and the feasibility of the developed control are also confirmed by simulations. Yongduan Song 0001, Junxia Guo, Xiucai Huang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | Observing the evolution of social network on Weibo by sampled dataabstractAlthough there have been many researches on the online social networks (OSNs), observing the evolution of a real OSN is still interesting and instructive for understanding people's behavior in OSNs. In this paper, the actual evolution of the social graph of a real OSN - Weibo, is studied by sampled data. The exact timestamp of creating or removing each following relationship cannot be sampled. However, by the created time of the users' accounts, the evolution of the social network of Weibo is roughly observed. In this way, it is found that the growing pattern of the network scale shows S-shape. Some other properties of the network, such as network density, the number of connected components, the efficiency of the network, clustering coefficient, degree assortativity, and so on, are also observed. As the network grows, the density of the network keeps reducing and eventually reaches a steady state. The change of the number of connected components indicates the users' crowd behavior during the network evolution. Junxia Guo |
ICIS | 3 |
| 2016 | Search Based Test Suite Minimization for Fault Detection and Localization: A Co-driven Method
Jingyao Geng, Zheng Li 0002, Ruilian Zhao, Junxia Guo |
SSBSE | 4 |
| 2013 | Analysis and Design of programmatic Interfaces for Integration of Diverse Web ContentsabstractThe technology that integrates various types of Web contents to build a new Web application through end-user programming is widely used nowadays. However, the Web contents do not have a uniform interface for accessing the data and computation. Most of the general Web users access information on the Web through applications until now. Hence, designing a uniform and flexible programmatic interface for integration of different Web contents is unavoidable. In this paper, we propose an approach that can be used to analyze Web applications automatically and reuse the information of Web applications through the programmatic interface we designed. Our approach can support the flexible integration of Web applications, Web services and Web feeds. In our experiments, we use a large number of Web pages from different types of Web applications and achieve the integration by the proposed programmatic interfaces. The experimental results show that our approach brings to the end-users a flexible and user-friendly programming environment. Junxia Guo |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2010 | Towards Flexible Mashup of Web Applications Based on Information Extraction and Transfer
Junxia Guo, Takehiro Tokuda |
WISE | 1 |
| 2010 | Deep mashup: a description-based framework for lightweight integration of web contentsabstractIn this paper, we present a description-based mashup framework for lightweight integration of Web contents. Our implementation shows that we can integrate not only the Web services but also the Web applications easily, even the Web contents dynamically generated by client-side scripts. Junxia Guo, Takehiro Tokuda |
WWW | 2 |
| 2009 | A New Partial Information Extraction Method for Personal Mashup ConstructionabstractNowadays more and more Web sites generate Web pages containing client-side scripts such as JavaScript and Flash instead of ordinary static HTML pages. These scripts create dynamic HTML pages and provide modern interfaces to Web applications such as AJAX applications. Unfortunately, for partial information extraction of Web pages, existing methods cannot extract dynamic Web contents. Junxia Guo, Takehiro Tokuda |
EJC | 1 |