VLDB 2026 Research / reviewers in the wild / expert
Lei Xu 0003
dblp:19/360-3
· DBLP profile ↗
39ranked-venue papers
7as first author
10since 2021 · last 2024
0000-0002-4815-2850ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 24 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorHuman-computer interaction and ubiquitous computing · 4 · 3 first-authorSystems, architecture and hardware · 3 · 1 first-author · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Python meets JIT compilers: A simple implementation and a comparative evaluationabstractAbstract Developing a just‐in‐time (JIT) compiler can be a daunting task, especially for a language as flexible as Python. While PyPy, powered with JIT compilation, can often outperform the official pure interpreter, CPython, by a noteworthy margin, its popularity remains far from comparable to that of CPython due to some issues. Given that an easier‐to‐deploy and better‐compatible JIT compiler would benefit more Python users, we have developed comPyler, a simple JIT compiler functioning as a CPython extension and intended to convert frequently interpreted CPython bytecode into equivalent machine code. Designed with good compatibility in mind, it does not alter CPython's internal data structures or external interfaces. Based on LLVM's mature infrastructure, it can be readily ported to almost all platforms. Compared with CPython, it achieved the highest speedup of 2.205, with an average of 1.093. Despite its relatively limited effect, comPyler incurs low development costs. As a baseline compiler, it also sheds light on the improvement attainable by optimizing solely the overhead of bytecode interpretation. Furthermore, as there is still a dearth of empirical research covering the multitude of JIT compilers available for Python, we have conducted a performance study that examines Jython, IronPython, PyPy, GraalPy, Pyston, Pyjion, and our comPyler. Our research takes into account not only the benchmark speed for various time windows but also the boot latency and memory footprint. Through this comprehensive study, our objective is to assist developers in gaining a better understanding of the effects of distinct JIT compilation techniques and to aid users in making informed decisions when choosing among different Python implementations. Qiang Zhang 0046, Lei Xu 0003, Baowen Xu |
Softw. Pract. Exp. | 2 |
| 2023 | NodeRT: Detecting Races in Node.js Applications PracticallyabstractNode.js has become one of the most popular development platforms due to its superior concurrency support. However, races induced by the nondeterministic execution order of event handlers may occur in Node.js applications, causing serious runtime failures. The state-of-the-art Node.js race detector NRace builds a happens-before (HB) graph before detection with a set of HB relation rules. In detection, NRace utilizes a heavy-weight BFS-based algorithm to query the reachability between resource operations, which introduces substantial overhead in practice, causing NRace inapplicable to real-world Node.js application test processes. This paper proposes a more practical Node.js dynamic race detection approach called NodeRT (Node.js Race Tracker). To reduce unnecessary overhead, NodeRT simplifies the HB relation rules, and divides the detection into three stages: trace collection stage, race candidate detection stage, and false positive removal stage. In the trace collection stage, NodeRT constructs a partial HB graph called asynchronous call tree (ACTree), enabling efficient reachability queries between event handlers. In the race candidate detection stage, NodeRT performs detection on the ACTree, which effectively eliminates most non-racing event handlers and outputs race candidates. In the false positive removal stage, NodeRT utilizes matching rules derived from HB relation rules and features of resources to reduce false positives in the race candidates. In experiments, NodeRT detects all known races and 9 unknown harmful races in real-world applications, whereas NRace only detects 3 of the unknown harmful races, with 64× more time consumption on average. Compared with NRace, NodeRT significantly reduces the overhead, making it practical to be integrated into real-world test processes. Jingyao Zhou, Lei Xu 0003, Gongzheng Lu, Weifeng Zhang 0001, Xiangyu Zhang 0001 |
ISSTA | 2 |
| 2023 | How Dynamic Features Affect API Usages? An Empirical Study of API Misuses in Python ProgramsabstractIncorrect usages of Application Programming Interfaces (APIs) may lead to unexpected problems during the software development process. Although there have been many attempts to address API-misuse issues, most of them are mainly for static languages. In contrast, API misuses in dynamic languages are rarely covered, mostly due to challenges about dynamic features. In this paper, we develop the first-ever comprehensive study of API misuses for Python programs. To accomplish this, we manually analyze 79,096 commits of six popular open-source Python projects on GitHub to collect true-positive cases. Based on the validation, we develop a classification of Python API Misuses, called PAM, and a dataset, PAMBench, containing 670 validated real-world API-misuse cases in popular Python programs. For each API-misuse case, we explore its root cause, symptom, program issue and repair method. Specifically, we pay attention to the effect of dynamic features on API usages in Python. The systematic study on PAMBench shows that, most importantly, dynamic features, especially type dynamics, have a non-negligible impact on API usages in Python, mainly related to incorrect assumptions about the type, callable state, attribute and existence of caller object, method call itself, passed argument(s) and return value during an API invocation. Our root-cause analysis reveals the importance of correct design, implementation, annotation, checking and recording about the types and states of all parts of API method calls during Python program development. Finally, we present possible solutions for more secure, reliable and maintainable API usages in Python. Xincheng He, Xiaojin Liu 0005, Lei Xu 0003, Baowen Xu |
SANER | 3 |
| 2023 | Python API Misuse Mining and Classification Based on Hybrid Analysis and Attention MechanismabstractAPIs play a crucial role in contemporary software development, streamlining implementation and maintenance processes. However, improper API usage can result in significant issues such as unexpected outcomes, security vulnerabilities and system crashes. To detect API misuses, current methods primarily rely on comparing established API usage patterns with target points for automated detection, mainly based on pre-validated datasets. Nonetheless, there is a scarcity of publicly available datasets on API misuses and their corresponding fixes, which hinders data-driven research. Moreover, most existing techniques concentrate on statically typed languages, such as Java and C, with only a few addressing dynamic languages like Python effectively, due to difficulties in handling dynamic features. Therefore, it is essential to identify Python API misuses and their fixes automatically and promptly. In this paper, we introduce HatPAM, a Hybrid Analysis and Attention-based Python API-Misuse Miner, which (a) provides a method for automatically mining true-positive commits related to Python API-misuse fixes from GitHub and (b) presents the subsequent processing for classifying Python API misuses in true-positive cases. Particularly, HatPAM applies hybrid static analysis and introduces a structure-based attention mechanism to examine syntax, semantics and structural features in Python code context, and considers the consistency between code and developers’ natural intent to significantly reduce false-positive cases. Evaluation on six popular Python projects reveals that HatPAM outperforms various state-of-the-art baselines, achieving up to 92.2% Precision, 86.7% Recall and 89.3% F1-score, indicating its capability to identify and classify Python API-misuse commits. Xincheng He, Xiaojin Liu 0005, Lei Xu 0003 |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2023 | RegCPython: A Register-based Python Interpreter for Better PerformanceabstractInterpreters are widely used in the implementation of many programming languages, such as Python, Perl, and Java. Even though various JIT compilers emerge in an endless stream, interpretation efficiency still plays a critical role in program performance. Does a stack-based interpreter or a register-based interpreter perform better? The pros and cons of the pair of architectures have long been discussed. The stack architecture is attractive for its concise model and compact bytecode, but our study finds that the register-based interpreter can also be implemented easily and that its bytecode size only grows by a small margin. Moreover, the latter turns out to be appreciably faster. Specifically, we implemented an open source Python interpreter named RegCPython based on CPython v3.10.1. The former is register based, while the latter is stack based. Without changes in syntax, Application Programming Interface, and Application Binary Interface, RegCPython is excellently compatible with CPython, as it does not break existing syntax or interfaces. It achieves a speedup of 1.287 on the most favorable benchmark and 0.977 even on the most unfavorable benchmark. For all Python-intensive benchmarks, the average speedup reaches 1.120 on x86 and 1.130 on ARM. Our evaluation work, which also serves as an empirical study, provides a detailed performance survey of both interpreters on modern hardware. It points out that the register-based interpreters are more efficient mainly due to the elimination of machine instructions needed, while changes in branch mispredictions and cache misses have a limited impact on performance. Additionally, it confirms that the register-based implementation is also satisfactory in terms of memory footprint, compilation cost, and implementation complexity. Qiang Zhang 0046, Lei Xu 0003, Baowen Xu |
ACM Trans. Archit. Code Optim. | 2 |
| 2022 | HatCUP: hybrid analysis and attention based just-in-time comment updatingabstractWhen changing code, developers sometimes neglect updating the related comments, bringing inconsistent or outdated comments. These comments increase the cost of program understanding and greatly reduce software maintainability. Researchers have put forward some solutions, such as CUP and HEBCUP, which update comments efficiently for simple code changes (i.e. modifying of a single token), but not good enough for complex ones. In this paper, we propose an approach named HatCUP (Hybrid Analysis and Attention based Comment UPdater), to provide a new mechanism for comment updating task. HatCUP pays attention to hybrid analysis and information. First, HatCUP considers the code structure change information and introduces a structure-guided attention mechanism combined with code change graph analysis and optimistic data flow dependency analysis. With a generally popular RNN-based encoder-decoder architecture, HatCUP takes the action of the code edits, the syntax, semantics and structure code changes, and old comments as inputs and generates a structural representation of the changes in the current code snippet. Furthermore, instead of directly generating new comments, HatCUP proposes a new edit or non-edit mechanism to mimic human editing behavior, by generating a sequence of edit actions and constructing a modified RNN model to integrate newly developed components. Evaluation on a popular dataset demonstrates that HatCUP outperforms the state-of-the-art deep learning-based approaches (CUP) by 53.8% for accuracy, 31.3% for recall and 14.3% for METEOR of the original metrics. Compared with the heuristic-based approach (HEBCUP), HatCUP also shows better overall performance. Hongquan Zhu, Xincheng He, Lei Xu 0003 |
ICPC | 3 |
| 2022 | An Empirical Study on the Impact of Python Dynamic Typing on the Project MaintenanceabstractPython is a popular typical dynamic programming language. In Python, dynamic typing is one of the most critical dynamic features. The lack of type information is likely to hinder the maintenance of Python projects. However, existing work has seldom focused on studying the impact of Python dynamic typing on project maintenance. This paper focuses on the two most common practices of Python dynamic typing, i.e. inconsistent-type assignments (ITA) and inconsistent variable types (IVT). Two approaches are proposed to identify ITA and IVT, i.e. identifying ITA by analyzing Abstract Syntax Trees and comparing identifiers types and identifying IVT by constructing a type dependency graph. In empirical experiments, we first locate the usage of ITA and IVT in 10 open-source Python projects. Then, we investigate the relations between the occurrence of ITA and IVT and the results of maintenance tasks. The study results show that projects are more prone to change as the number of dynamic typing identifiers increases. There is a weak connection between change-proneness and variable dynamic typing. There is a high probability that maintenance time and the acceptance of commits decrease as dynamic typing identifiers increase in projects. These results implicate that dynamic and static variables should be divided while developing new programming languages. Dynamic typing identifiers may not be the direct root causes for most software bugs. The categories of these bugs are worth exploring. Xinmeng Xia, Yanyan Yan, Xincheng He, Di Wu 0014, Lei Xu 0003, Baowen Xu |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2022 | Quantifying the interpretation overhead of Python
Qiang Zhang 0046, Lei Xu 0003, Xiangyu Zhang 0001, Baowen Xu |
Sci. Comput. Program. | 2 |
| 2021 | PyART: Python API Recommendation in Real-TimeabstractAPI recommendation in real-time is challenging for dynamic languages like Python. Many existing API recommendation techniques are highly effective, but they mainly support static languages. A few Python IDEs provide API recommendation functionalities based on type inference and training on a large corpus of Python libraries and third-party libraries. As such, they may fail to recommend or make poor recommendations when type information is missing or target APIs are project-specific. In this paper, we propose a novel approach, PyART, to recommend APIs for Python programs in real-time. It features a light-weight analysis to derives so-called optimistic data-flow, which is neither sound nor complete, but simulates the local data-flow information humans can derive. It extracts three kinds of features: data-flow, token similarity, and token co-occurrence, in the context of the program point where a recommendation is solicited. A predictive model is trained on these features using the Random Forest algorithm. Evaluation on 8 popular Python projects demonstrates that PyART can provide effective API recommendations. When historic commits can be leveraged, which is the target scenario of a state-of-the-art tool ARIREC, our average top-1 accuracy is over 50% and average top-10 accuracy over 70%, outperforming APIREC and Intellicode (i.e., the recommendation component in Visual Studio) by 28.48%-39.05% for top-1 accuracy and 24.41%-30.49% for top-10 accuracy. In other applications such as when historic comments are not available and cross-project recommendation, PyART also shows better overall performance. The time to make a recommendation is less than a second on average, satisfying the real-time requirement. Xincheng He, Lei Xu 0003, Xiangyu Zhang 0001, Yang Feng 0003, Baowen Xu |
ICSE | 2 |
| 2021 | Prioritizing code documentation effort: Can we do it simpler but better?
Shiran Liu, Zhaoqiang Guo, Yanhui Li 0001, Hongmin Lu, Lin Chen 0015, Lei Xu 0003, Yuming Zhou, Baowen Xu |
Inf. Softw. Technol. | 6 |
| 2020 | Composite Backdoor Attack for Deep Neural Network by Mixing Existing Benign FeaturesabstractWith the prevalent use of Deep Neural Networks (DNNs) in many applications, security of these networks is of importance. Pre-trained DNNs may contain backdoors that are injected through poisoned training. These trojaned models perform well when regular inputs are provided, but misclassify to a target output label when the input is stamped with a unique pattern called trojan trigger. Recently various backdoor detection and mitigation systems for DNN based AI applications have been proposed. However, many of them are limited to trojan attacks that require a specific patch trigger. In this paper, we introduce composite attack, a more flexible and stealthy trojan attack that eludes backdoor scanners using trojan triggers composed from existing benign features of multiple labels. We show that a neural network with a composed backdoor can achieve accuracy comparable to its original version on benign data and misclassifies when the composite trigger is present in the input. Our experiments on 7 different tasks show that this attack poses a severe threat. We evaluate our attack with two state-of-the-art backdoor scanners. The results show none of the injected backdoors can be detected by either scanner. We also study in details why the scanners are not effective. In the end, we discuss the essence of our attack and propose possible defense. Lei Xu 0003, Yingqi Liu, Xiangyu Zhang 0001 |
CCS | 2 |
| 2020 | CPC: automatically classifying and propagating natural language comments via program analysisabstractCode comments provide abundant information that have been leveraged to help perform various software engineering tasks, such as bug detection, specification inference, and code synthesis. However, developers are less motivated to write and update comments, making it infeasible and error-prone to leverage comments to facilitate software engineering tasks. In this paper, we propose to leverage program analysis to systematically derive, refine, and propagate comments. For example, by propagation via program analysis, comments can be passed on to code entities that are not commented such that code bugs can be detected leveraging the propagated comments. Developers usually comment on different aspects of code elements like methods, and use comments to describe various contents, such as functionalities and properties. To more effectively utilize comments, a fine-grained and elaborated taxonomy of comments and a reliable classifier to automatically categorize a comment are needed. In this paper, we build a comprehensive taxonomy and propose using program analysis to propagate comments. We develop a prototype CPC, and evaluate it on 5 projects. The evaluation results demonstrate 41573 new comments can be derived by propagation from other code locations with 88% accuracy. Among them, we can derive precise functional comments for 87 native methods that have neither existing comments nor source code. Leveraging the propagated comments, we detect 37 new bugs in open source large projects, 30 of which have been confirmed and fixed by developers, and 304 defects in existing comments (by looking at inconsistencies between existing and propagated comments), including 12 incomplete comments and 292 wrong comments. This demonstrates the effectiveness of our approach. Our user study confirms propagated comments align well with existing comments in terms of quality. Juan Zhai, Xiangzhe Xu, Guanhong Tao 0001, Minxue Pan, Shiqing Ma, Lei Xu 0003, Weifeng Zhang 0001, Lin Tan 0001, Xiangyu Zhang 0001 |
ICSE | 7 |
| 2020 | Correlations between deep neural network model coverage criteria and model qualityabstractInspired by the great success of using code coverage as guidance in software testing, a lot of neural network coverage criteria have been proposed to guide testing of neural network models (e.g., model accuracy under adversarial attacks). However, while the monotonic relation between code coverage and software quality has been supported by many seminal studies in software engineering, it remains largely unclear whether similar monotonicity exists between neural network model coverage and model quality. This paper sets out to answer this question. Specifically, this paper studies the correlation between DNN model quality and coverage criteria, effects of coverage guided adversarial example generation compared with gradient decent based methods, effectiveness of coverage based retraining compared with existing adversarial training, and the internal relationships among coverage criteria. Shenao Yan, Guanhong Tao 0001, Xuwei Liu, Juan Zhai, Shiqing Ma, Lei Xu 0003, Xiangyu Zhang 0001 |
ESEC/SIGSOFT FSE | 6 |
| 2020 | Black-box adversarial sample generation based on differential evolution
Lei Xu 0003, Yingqi Liu, Xiangyu Zhang 0001 |
J. Syst. Softw. | 2 |
| 2019 | Transferring Java Comments Based on Program Static Analysis
Binger Li, Xinlin Huang, Xincheng He, Lei Xu 0003 |
WISA | 5 |
| 2019 | Mining Core Contributors in Open-Source Projects
Xiaojin Liu 0005, Jiayang Bai, Lanfeng Liu, Hongrong Ouyang, Lei Xu 0003 |
WISA | 6 |
| 2019 | Grading Programs Based on Hybrid Analysis
Lei Xu 0003 |
WISA | 2 |
| 2019 | Semantic Web Service Discovery Based on LDA Clustering
Lei Xu 0003 |
WISA | 3 |
| 2019 | Hunting for bugs in code coverage tools via randomized differential testingabstractReliable code coverage tools are critically important as it is heavily used to facilitate many quality assurance activities, such as software testing, fuzzing, and debugging. However, little attention has been devoted to assessing the reliability of code coverage tools. In this study, we propose a randomized differential testing approach to hunting for bugs in the most widely used C code coverage tools. Specifically, by generating random input programs, our approach seeks for inconsistencies in code coverage reports produced by different code coverage tools, and then identifies inconsistencies as potential code coverage bugs. To effectively report code coverage bugs, we addressed three specific challenges: (1) How to filter out duplicate test programs as many of them triggering the same bugs in code coverage tools; (2) how to automatically reduce large test programs to much smaller ones that have the same properties; and (3) how to determine which code coverage tools have bugs? The extensive evaluations validate the effectiveness of our approach, resulting in 42 and 28 confirmed/fixed bugs for gcov and llvm-cov, respectively. This case study indicates that code coverage tools are not as reliable as it might have been envisaged. It not only demonstrates the effectiveness of our approach, but also highlights the need to continue improving the reliability of code coverage tools. This work opens up a new direction in code coverage validation which calls for more attention in this area. Yibiao Yang, Yuming Zhou, Hao Sun 0021, Zhendong Su 0001, Zhiqiang Zuo 0002, Lei Xu 0003, Baowen Xu |
ICSE | 6 |
| 2019 | Predictive analysis for race detection in software-defined networks
Gongzheng Lu, Lei Xu 0003, Yibiao Yang, Baowen Xu |
Sci. China Inf. Sci. | 2 |
| 2018 | Classifying Python Code Comments Based on Supervised Learning
Lei Xu 0003, Yanhui Li 0001 |
WISA | 2 |
| 2018 | Malicious JavaScript Code Detection Based on Hybrid AnalysisabstractJavaScript plays an important role in web applications and services, which is used by millions of web pages in optimizing interface design, embedding dynamic texts, reading and writing HTML elements, validating form data, responding to browser events, controlling cookies and much more. However, since JavaScript is cross-platform and can be executed dynamically, it has been a major vehicle for web-based attacks. Existing solutions work by performing static analysis or monitoring program execution dynamically. However, since the heavy use of obfuscation techniques, many methods no longer apply to malicious JavaScript code detection, and it has been a huge challenge to de-obfuscate obfuscated malicious JavaScript code accurately. In this paper, we propose a hybrid analysis method combining static and dynamic analysis for detecting malicious JavaScript code that works by first conducting syntax analysis and dynamic instrumentation to extract internal features that are related to malicious code and then performing classification-based detection to distinguish attacks. In addition, based on code instrumentation, we propose a new method which can deobfuscate part of obfuscated malicious JavaScript code accurately. Ultimately, we implement a browser plug-in called MJDetector and perform evaluation on 450 real web pages. Evaluation results show that our method can detect malicious JavaScript code and de-obfuscate obfucation effectively and efficiently. Specifically, MJDetector can detect JavaScipt attacks in current web pages with high accuracy 94.76% and de-obfuscate obfuscate code of specific types with accuracy 100% whereas the base line method can only detect with accuracy 81.16% and has no capacity of de-obfuscation. Xincheng He, Lei Xu 0003, Chunliu Cha |
APSEC | 2 |
| 2018 | A Caching-Based Parallel FP-Growth in Apache Spark
Zhicheng Cai, Xingyu Zhu 0005, Yuehui Zheng, Duan Liu, Lei Xu 0003 |
ICA3PP (3) | 5 |
| 2016 | An Empirical Study on the Characteristics of Python Fine-Grained Source Code Change TypesabstractSoftware has been changing during its whole life cycle. Therefore, identification of source code changes becomes a key issue in software evolution analysis. However, few current change analysis research focus on dynamic language software. In this paper, we pay attention to the fine-grained source code changes of Python software. We implement an automatic tool named PyCT to extract 77 kinds of fine-grained source code change types from commit history information. We conduct an empirical study on ten popular Python projects from five domains, with 132294 commits, to investigate the characteristics of dynamic software source code changes. Analyzing the source code changes in four aspects, we distill 11 findings, which are summarized into two insights on software evolution: change prediction and fault code fix. In addition, we provide direct evidence on how developers use and change dynamic features. Our results provide useful guidance and insights for improving the understanding of source code evolution of dynamic language software. Zhifei Chen, Wanwangying Ma, Lin Chen 0015, Lei Xu 0003, Baowen Xu |
ICSME | 5 |
| 2016 | Statically Detect Data Races for WS-BPEL Web Services by Constraint SolverabstractNowadays, Web services are widely used because of their interoperability and reusability. Multiple Web services can be composed following some business logic specified by BPEL (Business Process Execution Language) scripts. Since BPEL scripts allow specifying concurrent workflow, typical concurrency problems, such as data race, atomicity violation and order violation, also commonly occur in BPEL scripts. These issues are hard to detect and reproduce due to their non-determinism and the special language features of BPEL. In this paper, we implement a tool to detect data races for WS-BPEL based on static analysis approach and constraints solver. Our system is based on three key concepts: (1) a preprocess model to record necessary information, (2) a thorough Happens-Before model of WS-BPEL concurrency, (3) constraint encoding to transfer Happens-Before relationship to constraints and check if there is a feasible solution (namely data races) by Z3-Str solver. We evaluate the usability and performance of our tool on 10 benchmark programs with effective results. Lei Xu 0003, Baowen Xu, Weifeng Zhang 0001 |
ICWS | 1 |
| 2016 | ARROW: automated repair of races on client-side web pagesabstractModern browsers have a highly concurrent page rendering process in order to be more responsive. However, such a concurrent execution model leads to various race issues. In this paper, we present ARROW, a static technique that can automatically, safely, and cost effectively patch certain race issues on client side pages. It works by statically modeling a web page as a causal graph denoting happens-before relations between page elements, according to the rendering process in browsers. Races are detected by identifying inconsistencies between the graph and the dependence relations intended by the developer. Detected races are fixed by leveraging a constraint solver to add a set of edges with the minimum cost to the causal graph so that it is consistent with the intended dependences. The input page is then transformed to respect the repair edges. ARROW has fixed 151 races from 20 real world commercial web sites. Weihang Wang 0001, Yunhui Zheng, Peng Liu 0010, Lei Xu 0003, Xiangyu Zhang 0001, Patrick Eugster |
ISSTA | 4 |
| 2016 | Effort-aware just-in-time defect prediction: simple unsupervised models could be better than supervised modelsabstractUnsupervised models do not require the defect data to build the prediction models and hence incur a low building cost and gain a wide application range. Consequently, it would be more desirable for practitioners to apply unsupervised models in effort-aware just-in-time (JIT) defect prediction if they can predict defect-inducing changes well. However, little is currently known on their prediction effectiveness in this context. We aim to investigate the predictive power of simple unsupervised models in effort-aware JIT defect prediction, especially compared with the state-of-the-art supervised models in the recent literature. We first use the most commonly used change metrics to build simple unsupervised models. Then, we compare these unsupervised models with the state-of-the-art supervised models under cross-validation, time-wise-cross-validation, and across-project prediction settings to determine whether they are of practical value. The experimental results, from open-source software systems, show that many simple unsupervised models perform better than the state-of-the-art supervised models in effort-aware JIT defect prediction. Yibiao Yang, Yuming Zhou, Hongmin Lu, Lei Xu 0003, Baowen Xu, Hareton K. N. Leung |
SIGSOFT FSE | 6 |
| 2016 | Empirical analysis of network measures for predicting high severity software faults
Lin Chen 0015, Wanwangying Ma, Yuming Zhou, Lei Xu 0003, Ziyuan Wang 0001, Zhifei Chen, Baowen Xu |
Sci. China Inf. Sci. | 4 |
| 2015 | Generating Test Cases for Composite Web Services by Parsing XML Documents and Solving ConstraintsabstractWeb services are widely used nowadays for their interoperability and reusability. Since Web services only provide interface information for users and source codes are encapsulated, generating test cases for Web services in the view of users has more challenges than traditional software. We develop a constraint-solver based method to generate test cases for composite Web services. The technique first parses the related files, such as XSD (XML Schema Definition), WSDL (Web Service Description Language) and BPEL (Business Process Executing Language) scripts to obtain the constraints for variable types, in-out relations, conditions and orders. Then, by using the Z3-str solver, test cases are generated according to different testing coverage criterions. Our evaluation results indicate that our method is effective to generate test cases for Web services with high coverage and low redundancy. Lei Xu 0003, Baowen Xu |
COMPSAC | 2 |
| 2013 | Recommending Web Service Based on User Relationships and PreferencesabstractWith the popularity of social network and the increasing number of Web Services, making individual service recommendation has been a hot research spot nowadays. In this paper, we present a service recommendation algorithm named as URPC-Rec (User Relationships & Preferences Clustering and Recommendation), which first clusters users based on their history behaviors such as the services they ever invoked, and then makes personalized recommendations for users considering both the clustering results and user basic information and relationships, such as gender, age, occupation, preference tags, etc. The case study indicates that URPC-Rec can effectively reduce the dimensionality of sparse matrix, and partially solve the cold-start problem of recommendation systems. The comprehensive experiment shows that URPC-Rec algorithm with user relationships and references has better recommending result than the one without user information and the collaborative filtering approach. Zhaogui Xu, Lei Xu 0003, Yanhui Li 0001, Lin Chen 0015 |
ICWS | 3 |
| 2011 | Web Service Discovery Based on User RequirementsabstractWith the rapid development of Web services, Web service discovery becomes a significant challenge in the matching precision and efficiency. In this paper, we present a novel method in the view of users' requirements for Web Service discovery. We set up the requirement model firstly, so as to present users' requirements in details, then cluster the Web Services due to users' common requirements, and obtain corresponding Web services' QoS values. Thus, users can find their really needed Web services in an ordered list. The case study indicates the effect of our method, namely, in the view of users' requirements, we sort the candidate Services due to their functions and their QoS attributes in descending order. Lei Xu 0003, Baowen Xu, Lianjie Chen |
HPCC | 1 |
| 2008 | Combining MDE and UML to Reverse Engineer Web-Based Legacy SystemsabstractThe research in this paper focuses on an approach to reverse engineering Web-based legacy systems with the integration of model-driven engineering and UML. Three types of link-based models of Web-based legacy systems are presented. Web-based legacy systems are parsed to find judgement conditions of model, and UML diagrams are described based on the modelling rules. Jianjun Pu, Baowen Xu, Lei Xu 0003, William C. Chu |
COMPSAC | 4 |
| 2007 | Applying Agent into Intelligent Web Application TestingabstractWeb application testing is concerned with numerous and complicated testing objects, methods and processes. In order to improve the testing efficiency, we firstly analyze the necessity and feasibility of the automatic and intelligent testing for Web applications; Then, we discuss several scenes of applying agent into Web application testing, such as using agent to obtain users' visiting actions, carry out performance testing, regression testing and usability evolvement; next, we adopt agent to execute the testing, including the testing process and the detailed actions, so as to monitor, manage and handler exceptions during the whole testing execution. Thus, in this way, the Web application testing can be completed more automatically and intelligently. Lei Xu 0003, Baowen Xu |
CW | 1 |
| 2005 | Configuration Strategies for Evolutionary TestingabstractThis paper presents a new approach to generating configuration-oriented executable symbolic test sequences from extended finite state machine (EFSM) models. The information about the values of the context variables and the domain intervals of the input parameters are exploited to guide the derivation of the test sequences. Meanwhile, the transition guards along the test sequences are continually used to reduce the domain intervals of the input parameters. Experiments indicate that this method significantly reduces the EFSM state space to be explored and the number of non-executable symbolic test sequences to be generated. Since parameterized input events are allowed to occur in EFSM cycles, this method is suitable for testing the open reactive systems that interact with the environments via parameterized input events. Xiaoyuan Xie, Baowen Xu, Changhai Nie, Lei Xu 0003 |
COMPSAC (2) | 5 |
| 2005 | Research on the Analysis and Measurement for Testing Results of Web ApplicationsabstractReasonable analysis and corrective measurement for the testing results of Web applications can effectively judge the effect and efficiency of the testing. Therefore, based on the previous work, we propose a new method for testing results analysis and comparison, which uses the semantic label and XML description technique to realize the information separation between data and display in the Web pages, so as to directly compare the testing results and the expected results. Furthermore, combined with the realities, we determine the metric indexes of Web application testing, so as to provide the criterions and guidelines for the evaluations of the Web applications and their testing processes. And we introduce the feedback control mechanism into the development and evolvement of Web applications, so as to further improve the system quality. Lei Xu 0003, Baowen Xu, Yanxiang He, Hanwu Chen, Qiaoming Zhu |
CW | 1 |
| 2004 | A Framework for Web Applications TestingabstractWeb application testing is concerned with numerous and complicated testing objects, methods and processes. So a testing framework fitting for the properties of Web application is needed to guide and organize all the testing tasks. Based on the analysis for Web application characters and traditional software testing process, the process for Web application testing is modeled, which describes a series of testing flows such as the testing requirement analysis, test cases generation and selection, testing execution, and testing results analysis and measurement. Furthermore, the realization techniques are also investigated so as to integrate each testing step and implement the whole testing process harmoniously and effectively. Thus the framework is suitable for the Internet environment and can guide the Web application testing actively and availably. Lei Xu 0003, Baowen Xu |
CW | 1 |
| 2003 | Regression Testing for Web Applications Based on SlicingabstractWeb applications have rapid developing speed and changeable user demands, so the regression testing is much important. Since the changed demands result in different versions of Web applications, and the faults usually hiding in the adjusted contents, the regression testing must cover all the related pages. In order to carry through the regression testing quickly and effectively, we make the simplification with the method of slicing. Firstly, we analyze the possible changes in the Web applications and the influences produced by these changes, discussing in the direct-dependent and indirect-dependent way; next, we give the regression testing method based on slicing emphasized on the indirect-dependent among data, i.e., obtaining the dependent set of changed variables by forward and backward search method and generating the testing suits; conclusion remarks and future work are given at last. Lei Xu 0003, Baowen Xu, Zhenqiang Chen, Jixiang Jiang, Huowang Chen |
COMPSAC | 1 |
| 2003 | Parallel Algorithm for Mining Fuzzy Association RulesabstractThe principle and steps of the algorithm for mining fuzzy association rules is studied, and the parallel algorithm for mining fuzzy association rules is presented. In this parallel mining algorithm, quantitative attributes are partitioned into several fuzzy sets by the parallel fuzzy c-means algorithm, and fuzzy sets are applied to soften the partition boundary of the attributes. Then, the parallel algorithm for mining Boolean association rules is improved to discover frequent fuzzy attributes. Last, the fuzzy association rules with at least fuzzy confidence are generated on all processors. The parallel mining algorithm is implemented on the distributed linked PC/workstation. The experiment results show that the parallel mining algorithm has fine scaleup, sizeup and speedup. Baowen Xu, Jianjiang Lu, Yingzhou Zhang, Lei Xu 0003, Huowang Chen |
CW | 4 |
| 2003 | A Browser Compatibility Testing Method Based on Combinatorial Testing
Lei Xu 0003, Baowen Xu, Changhai Nie, Huowang Chen |
ICWE | 1 |