Baoquan Cui

dblp:244/6492 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
9since 2021 · last 2025
0009-0004-8218-1112ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 3 first-author · 9 since 2021
YearPublicationVenuePosition
2025 Static Analysis of Remote Procedure Call in Java Programs
abstract
The Remote Procedure Call (RPC) is commonly used for inter-process communications over network, allowing a program to invoke a procedure in another address space, even in another machine as if it were a local call. Its convenience comes from encapsulating network communication. However, for the same reason, it cannot be penetrated by current static analyzers. Since the RPC based programs/frameworks play a more important role in various domains, the static analysis of RPC is significant and cannot be ignored. We have observed that many of the existing RPC frameworks/programs written in Java are based on explicit protocols, which makes them possible to be modelled for static analysis. The challenges are how to identify RPC operations in different frameworks/programs and how to automatically establish relationships between clients and servers. In this paper, we propose a novel approach, RPCBridge, which uses an adapter to unify the most basic operations during the RPC process. It models the RPC with logic rules in a straightforward and precise way based on its semantics, performs points-to analysis and constructs RPC edges in the call graph, making it more complete. The evaluation on real-world large-scale Java programs based on 5 common RPC frameworks shows that our approach can effectively capture the operations of the RPC and construct critical links between clients and servers, in which 60.1 % are the true caller-callee pairs after execution. Our approach is expected to bring significant benefits (+24.3 % leakage paths for the taint analyzer) for previously incompletely modelled code with a very little memory and time overhead, and connect the modules in a system, so that it can be statically analyzed more holistically.
Baoquan Cui, Rong Qu
ICSE1
2025 An Empirical Study: Mems as a Static Performance Metric
abstract
Performance analysis is essential to ensure the non-functional performance requirements of a software system. However, existing runtime-based approaches suffer from the issues of efficiency and platform dependency. In this paper, we investigate the effectiveness of using the mems value to statically estimate the program performance. The mems value, originally proposed by Donald Knuth, is a static and architecture-independent metric used to measure memory access, to estimate program performance statically. We developed an instrumentation tool to record the control flow and measure the mems value by rewriting the source code. Experimental results across ten classical algorithm programs show that execution paths of a program with larger mems values consistently exhibit lower efficiency. Whereas the correlation weakens among different programs. This indicates that the mems maric is best suited for comparing the performance of various paths in the same program.
Baoquan Cui, Xutong Ma, Jian Zhang 0001
QRS2
2024 DMMPP: Constructing Dummy Main Methods for Android Apps with Path-Sensitive Predicates
abstract
Android is based on an event-driven model, which hides the main method, and is driven by the lifecycle methods and listeners from user interaction. FlowDroid, constructs a dummy main method statically emulating the lifecycle methods. The dummy main method has been widely used by FlowDroid and also other Android analyzers as their entry points. However, the existing dummy main method is not designed for path-sensitive analysis, whose paths may be unsatisfiable. Thus, when using original dummy main methods, path-sensitive analysis, e.g., symbolic execution, may suffer from infeasible paths. In this paper, we present DMMPP, the first dummy main method generator for Android applications with path-sensitive predicates, and the corresponding path condition is satisfiable. DMMPP constructs dummy main methods for the four types of components in an application with a more realistic simulation for the lifecycle methods. The experiment demonstrates the benefits of our tool for path-sensitive analyzers, improving 28.5 times more explored paths with a low time overhead.
Baoquan Cui, Jiwei Yan, Jian Zhang 0001
ISSTA1
2023 Detection of Java Basic Thread Misuses Based on Static Event Analysis
abstract
The fundamental asynchronous thread (java.lang. Thread) in Java can be easily misused, due to the lack of deep understanding for garbage collection and thread interruption mechanism. For example, a careless implementation of asynchronous thread may cause no response to the interrupt mechanism in time, resulting in unexpected thread-related behaviors, especially resource leak/waste. Currently, few works aim at these misuses and related works adopt either the dynamic approach which lacks effective inputs or the static path-sensitive approach with high time consumption due to the path explosion, causing false negatives. We have found that the behavior of threads and the interaction between threads and its referencing objects can be abstracted. In this paper, we propose an event analysis approach to detect the defects in Java programs and Android apps, which focuses on the existence or the order of the events to reduce the false negatives. We extract the misuse-related events, containing the thread events and the destroy events of the object referenced by the thread. Then we analyze the events with loop identification, happens-before relationship construction and alias determination. Finally, we implement an automatic tool named Leopard and evaluate it on real world Java programs and Android apps. Experiments show that it is efficient when comparing with the existing approach (misuse: 723 vs 47, time: 60s vs 30min), which also outperforms the existing work in precision. The manual check indicates that Leopard is more efficient and effective than existing work. Besides, 66 issues reported by us have been confirmed and 21 of them have been fixed by developers.
Baoquan Cui, Jiwei Yan, Jun Yan 0009, Jian Zhang 0001
ASE1
2023 PSMT: Satisfiability Modulo Theories Meets Probability Distribution
abstract
SMT (Satisfiability Modulo Theories) has been widely used in program verification, analysis, and test generation. But sometimes, SMT solver outputs incomprehensible solutions, especially for practical instances. Besides, due to the design of the deterministic algorithms, for a given formula, the result of each run is the same. In this paper, we concentrate on combining SMT solving with probability, which will instruct the SMT solver to give some plausible solutions. We define a special problem: PSMT, which allows solving an SMT instance with variables conforming to a certain distribution. We define distribution under constraint for PSMT, which is based on MCSAT (Model Constructing Satisfiability), a mainstream SMT-solving algorithm. We propose the Prob-MCSAT algorithm, which combines the MCSAT algorithm and introduces the probability to variables. The visualized examples show that the resulting assignments will form a clear trend based on Prob-SMT.
Fuqi Jia, Xutong Ma, Baoquan Cui, Minghao Liu 0001, Pei Huang 0002, Feifei Ma, Jian Zhang 0001
ASE4
2022 String Test Data Generation for Java Programs
abstract
Appropriate string test data generation is important for program testing. Complex string APIs combinations are commonly used to handle string parameters. However, the complex combinations make it difficult to express comprehensive string related constraints and generate suitable string data to trigger bugs and cover more branches. In this paper, we propose a novel approach to characterize the input strings and their operations (API invocations) with the regular expressions for string test data generation, with insight that they support rich syntax and can express the semantics of various string APIs combinations. We build a set of mapping rules that map 48 string APIs in Java to regular expressions, and design an inference algorithm to generate regular expressions for the complex string APIs combinations. With these regular expressions, more effective string data can be generated in an efficient way. Experiments on multi-type programs from assignments, LeetCode platform and open source community show that our approach can increase the branch coverage (17%) and find more bugs (+81) than the existing work. For the basic library JDK, 17 defects have been found, of which 14 are confirmed by the JDK developers and 3 are fixed in new version.
Baoquan Cui, Jiwei Yan, Jun Yan 0009, Jian Zhang 0001
ISSRE2
2022 ExcePy: A Python Benchmark for Bugs with Python Built-in Types
abstract
As bugs of Python built-in types can cause code crashes, detecting them is critical to the robustness of the software. Researchers have concluded plenty of patterns for the bug causes and applied these patterns in detection tools. But these tools are only evaluated on handcrafted bugs or bugs obtained from QA pages. Because such bugs cannot reflect the complex code structures and various bug types encountered in real-world projects, the evaluation result is untrustworthy when applied to these projects. As a result, a collection of real-world reproducible bugs is essential for tool evaluation and future bug-related research. In this paper, we propose ExcePy, a benchmark for providing bugs of Python built-in types. We collect 180 bugs from the evolution of 15 real-world open-source Python projects on GitHub and then manually build test scripts for bug reproduction. Meanwhile, to improve tool evaluation efficiency, we present a code pruning strategy that can minimize buggy code size while retaining bug reproducibility and apply it to ExcePy to provide simplified buggy code. To demonstrate the benefits of ExcePy, we use three static analyzers and two fuzzers to detect bugs collected in ExcePy. We found that simplified code can significantly reduce running time and avoid many tool crashes, and bugs supplied by ExcePy can reveal limitations of existing tools in reporting real-world bugs.
Rongjie Yan, Jiwei Yan, Baoquan Cui, Jun Yan 0009, Jian Zhang 0001
SANER4
2021 Dynamic Detection of AsyncTask Related Defects
abstract
As a widely used Android asynchronous component, AsyncTask is used to run time-consuming tasks. However, the misuse of AsyncTask will cause defects, i.e., crashes and memory leaks. Based on static analysis, existing approaches cannot accurately detect AsyncTask-related defects and produce many false positives since some paths are not reachable in practice. In this paper, we propose a dynamic detection method based on instrumentation, Monkey execution and log analysis to detect these defects. And we implement a tool AD2Checker based on the proposed method. Our experiment on 19 real-world apps shows that it has found 145 bugs and has no false positives. Moreover, it triggers crashes caused by misuse of AsyncTask.
Linjie Pan 0001, Baoquan Cui, Jun Yan 0009, Jian Zhang 0001
QRS3
2021 Are the Scala Checks Effective? Evaluating Checks with Real-world Projects
abstract
Static analyzers can assist developers in detecting flaws and improving software quality. An analyzer often has numerous checkers, each of which implements a different checking rule. These checks can create a lot of warnings in real-world projects, putting a lot of pressure on programmers to examine them. Thus, it is critical to assess the effectiveness of these checkers before putting them to use. Typically, time-consuming questionnaires or human assessments of the warnings are employed to evaluate the checkers, which results in inefficiency when applied to real-world work. The significance and accuracy of checkers are the topics of this research, with the first reflecting the developers' attention to the checkers and the second reflecting the false-positive rate. We focus on Scala checkers in particular because, despite the popularity of the Scala programming language, there has been little study on them. We propose a method for tracking warnings in real-world projects and assessing the two features for each checker. We use 115 checks and six well-known Scala apps to demonstrate our approach. Based on the 191k warnings delivered by these checkers, the approach can identify 154k false positives, and it finds that only around 1/5 of the checks can benefit developers.
Jiwei Yan, Baoquan Cui, Jun Yan 0009, Jian Zhang 0001
QRS3
2020 Static asynchronous component misuse detection for Android applications
abstract
Facing the limited resource of smartphones, asynchronous programming significantly improves the performance of Android applications. Android provides several packaged components to ease the development of asynchronous programming. Among them, the AsyncTask component is widely used by developers since it is easy to implement. However, the abuse of AsyncTask component can decrease responsiveness and even lead to crashes. By investigating the Android Developer Documentation and technical forums, we summarize five misuse patterns about AsyncTask. To detect them, we propose a flow, context, object and field-sensitive inter-procedural static analysis approach. Specifically, the static analysis includes typestate analysis, reference analysis and loop analysis. Based on the AsyncTask-related information obtained during static analysis, we check the misuse according to predefined detection rules. The proposed approach is implemented into a tool called AsyncChecker. We evaluate AsyncChecker on a self-designed benchmark suite called AsyncBench and 1,759 real-world apps. AsyncChecker finds 17,946 misused AsyncTask instances in 1,417 real-world apps (80.6%). The precision, recall and F-measure of AsyncChecker on real-world applications are 97.2%, 89.8% and 0.93, respectively. Compared with existing tools, AsyncChecker can detect more asynchronous problems. We report the misuse problems to developers via GitHub. Several developers have confirmed and fixed the problems found by AsyncChecker. The result implies that our approach is effective and developers do take the misuse of AsyncTask as a serious problem.
Linjie Pan 0001, Baoquan Cui, Jiwei Yan, Jun Yan 0009, Jian Zhang 0001
ESEC/SIGSOFT FSE2
2019 Androlic: an extensible flow, context, object, field, and path-sensitive static analysis framework for Android
abstract
Static analysis is widely used to detect potential defects in apps. Existing analysis tools focus on specific problems and vary in supported sensitivity, which make them difficult to reuse and extend for new analysis tasks. This paper presents Androlic, a precise static analysis framework for Android which is flow, context, object, field and path-sensitive. Through configuration items and APIs provided by Androlic, developers can easily extend it to perform custom analysis tasks. Evaluation on an example program and 20 real-world apps show that Androlic can analyze apps with high precision and efficiency.
Linjie Pan 0001, Baoquan Cui, Jiwei Yan, Xutong Ma, Jun Yan 0009, Jian Zhang 0001
ISSTA2